agora inbox for pgsql-hackers@postgresql.orghelp / color / mirror / Atom feed
[PATCH v6] README for 64bit xid 265+ messages / 2 participants [nested] [flat]
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v6] README for 64bit xid @ 2022-01-10 19:20 Pavel Borisov <pashkin.elfe@gmail.com> 0 siblings, 0 replies; 265+ messages in thread From: Pavel Borisov @ 2022-01-10 19:20 UTC (permalink / raw) Authors: - Pavel Borisov <pashkin.elfe@gmail.com> - Maxim Orlov <orlovmg@gmail.com> - Yura Sokolov <y.sokolov@postgrespro.ru> <funny.falcon@gmail.com> --- src/backend/access/heap/README.XID64 | 128 +++++++++++++++++++++++++++ 1 file changed, 128 insertions(+) create mode 100644 src/backend/access/heap/README.XID64 diff --git a/src/backend/access/heap/README.XID64 b/src/backend/access/heap/README.XID64 new file mode 100644 index 00000000000..457ba9b9ef5 --- /dev/null +++ b/src/backend/access/heap/README.XID64 @@ -0,0 +1,128 @@ +src/backend/access/heap/README.XID64 + +64-bit Transaction ID's (XID) +============================= + +A limited number (N = 2^32) of XID's required to do vacuum freeze to prevent +wraparound every N/2 transactions. This causes performance degradation due +to the need to exclusively lock tables while being vacuumed. In each +wraparound cycle, SLRU buffers are also being cut. + +With 64-bit XID's wraparound is effectively postponed to a very distant +future. Even in highly loaded systems that had 2^32 transactions per day +it will take huge 2^31 days before the first enforced "vacuum to prevent +wraparound"). Buffers cutting and routine vacuum are not enforced, and DBA +can plan them independently at the time with the least system load and least +critical for database performance. Also, it can be done less frequently +(several times a year vs every several days) on systems with transaction rates +similar to those mentioned above. + +On-disk tuple and page format +----------------------------- + +On-disk tuple format remains unchanged. 32-bit t_xmin and t_xmax store the +lower parts of 64-bit XMIN and XMAX values. Each heap page has additional +64-bit pd_xid_base and pd_multi_base which are common for all tuples on a page. +They are placed into a pd_special area - 16 bytes in the end of a heap page. +Actual XMIN/XMAX for a tuple are calculated upon reading a tuple from a page +as follows: + +XMIN = t_xmin + pd_xid_base. (1) +XMAX = t_xmax + pd_xid_base/pd_multi_base. (2) + +"Double XMAX" page format +--------------------------------- + +At first read of a heap page after pg_upgrade from 32-bit XID PostgreSQL +version pd_special area with a size of 16 bytes should be added to a page. +Though a page may not have space for this. Then it can be converted to a +temporary format called "double XMAX". + +All tuples after pg-upgrade would necessarily have xmin = FrozenTransactionId. +So we don't need tuple header t_xmin field and we reuse t_xmin to store higher +32 bits of its XMAX. + +Double XMAX format is only for full pages that don't have 16 bytes for +pd_special. So it neither has a place for a single tuple. Insert and HOT update +for double XMAX pages is impossible and not supported. We can only read or +delete tuples from it. + +When we are able to prune page double XMAX it will be converted from it to +general 64-bit XID page format with all operations on its tuples supported. + +In-memory tuple format +---------------------- + +In-memory tuple representation consists of two parts: +- HeapTupleHeader from disk page (contains all heap tuple contents, not only +header) +- HeapTuple with additional in-memory fields + +HeapTuple for each tuple in memory stores t_xid_base/t_multi_base - a copies of +page's pd_xid_base/pd_multi_base. With tuple's 32-bit t_xmin and t_xmax from +HeapTupleHeader they are used to calculate actual 64-bit XMIN and XMAX: + +XMIN = t_xmin + t_xid_base. (3) +XMAX = t_xmax + t_xid_base/t_multi_base. (4) + +The downside of this is that we can not use tuple's XMIN and XMAX right away. +We often need to re-read t_xmin and t_xmax - which could actually be pointers +into a page in shared buffers and therefore they could be updated by any other +backend. + +Update/delete with 64-bit XIDs and 32-bit t_xmin/t_xmax +-------------------------------------------------------------- + +When we try to delete/update a tuple, we check that XMAX for a page fits (2). +I.e. that t_xmax will not be over MaxShortTransactionId relative to +pd_xid_base/pd_multi_base of a its page. + +If the current XID doesn't fit a range +(pd_xid_base, pd_xid_base + MaxShortTransactionId) (5): + +- heap_page_prepare_for_xid() will try to increase pd_xid_base/pd_multi_base on +a page and update all t_xmin/t_xmax of the other tuples on the page to +correspond new pd_xid_base/pd_multi_base. + +- If it was impossible, it will try to prune and freeze tuples on a page. + +- If this is unsuccessful it will throw an error. Normally this is very +unlikely but if there is a very old living transaction with an age of around +2^32 this can arise. Basically, this is a behavior similar to one during the +vacuum to prevent wraparound when XID was 32-bit. Dba should take care and +avoid very-long-living transactions with an age close to 2^32. So long-living +transactions often they are most likely defunct. + +Insert with 64-bit XIDs and 32-bit t_xmin/t_xmax +------------------------------------------------ + +On insert we check if current XID fits a range (5). Otherwise: + +- heap_page_prepare_for_xid() will try to increase pd_xid_base for t_xmin will +not be over MaxShortTransactionId. + +- If it is impossible, then it will try to prune and freeze tuples on a page. + +Known issue: if pd_xid_base could not be shifted to accommodate a tuple being +inserted due to a very long-running transaction, we just throw an error. We +neither try to insert a tuple into another page nor mark the current page as +full. So, in this (unlikely) case we will get regular insert errors on the next +tries to insert to the page 'locked' by this very long-running transaction. + +Upgrade from 32-bit XID versions +-------------------------------- + +pg_upgrade doesn't change pages format itself. It is done lazily after. + +1. At first heap page read, tuples on a page are repacked to free 16 bytes +at the end of a page, possibly freeing space from dead tuples. + +2A. 16 bytes of pd_special is added if there is a place for it + +2B. Page is converted to "Double XMAX" format if there is no place for +pd_special + +3. If a page is in double XMAX format after its first read, and vacuum (or +micro-vacuum at select query) could prune some tuples and free space for +pd_special, prune_page will add pd_special and convert page from double XMAX +to general 64-bit XID page format. -- 2.24.3 (Apple Git-128) --cpok4wp6gsarlzvp-- ^ permalink raw reply [nested|flat] 265+ messages in thread
* [PATCH v52 06/10] Error out any process that would block at REPACK @ 2026-04-01 15:35 Antonin Houska <ah@cybertec.at> 0 siblings, 0 replies; 265+ messages in thread From: Antonin Houska @ 2026-04-01 15:35 UTC (permalink / raw) Any process waiting on REPACK to release its lock would actually cause it to deadlock when it tries to upgrade its lock to AEL, losing all work done to that point. We avoid this by teaching the deadlock detector to raise an error when this condition is detected. --- src/backend/commands/repack.c | 60 ++++++++++---- src/backend/storage/lmgr/deadlock.c | 15 ++++ src/include/storage/proc.h | 6 +- src/test/modules/injection_points/Makefile | 1 + .../expected/repack_deadlock.out | 63 ++++++++++++++ src/test/modules/injection_points/meson.build | 1 + .../specs/repack_deadlock.spec | 83 +++++++++++++++++++ 7 files changed, 210 insertions(+), 19 deletions(-) create mode 100644 src/test/modules/injection_points/expected/repack_deadlock.out create mode 100644 src/test/modules/injection_points/specs/repack_deadlock.spec diff --git a/src/backend/commands/repack.c b/src/backend/commands/repack.c index 03829892d57..d6e446d582d 100644 --- a/src/backend/commands/repack.c +++ b/src/backend/commands/repack.c @@ -279,6 +279,21 @@ ExecRepack(ParseState *pstate, RepackStmt *stmt, bool isTopLevel) /* Determine the lock mode to use. */ lockmode = RepackLockLevel((params.options & CLUOPT_CONCURRENT) != 0); + /* + * If in concurrent mode, set the PROC_IN_CONCURRENT_REPACK flag. This + * makes the deadlock checker cause anyone that would conflict with us to + * error out. It's important to set this flag ahead of actually locking + * the relation; it won't of course affect anyone until we do have a lock + * that others can conflict with. + */ + if ((params.options & CLUOPT_CONCURRENT) != 0) + { + LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE); + MyProc->statusFlags |= PROC_IN_CONCURRENT_REPACK; + ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags; + LWLockRelease(ProcArrayLock); + } + /* * If a single relation is specified, process it and we're done ... unless * the relation is a partitioned table, in which case we fall through. @@ -479,11 +494,8 @@ RepackLockLevel(bool concurrent) * If indexOid is InvalidOid, the table will be rewritten in physical order * instead of index order. * - * Note that, in the concurrent case, the function releases the lock at some - * point, in order to get AccessExclusiveLock for the final steps (i.e. to - * swap the relation files). To make things simpler, the caller should expect - * OldHeap to be closed on return, regardless CLUOPT_CONCURRENT. (The - * AccessExclusiveLock is kept till the end of the transaction.) + * On return, OldHeap is closed but locked with AccessExclusiveLock - the lock + * will be released at end of the transaction. * * 'cmd' indicates which command is being executed, to be used for error * messages. @@ -515,10 +527,12 @@ cluster_rel(RepackCommand cmd, Relation OldHeap, Oid indexOid, /* * Make sure we're not in a transaction block. * - * The reason is that repack_setup_logical_decoding() could deadlock - * if there's an XID already assigned. It would be possible to run in - * a transaction block if we had no XID, but this restriction is - * simpler for users to understand and we don't lose anything. + * The reason is that repack_setup_logical_decoding() could wait + * indefinitely for our XID to complete. (The deadlock detector would + * not recognize it because we'd be waiting for ourselves, i.e. no + * real lock conflict.) It would be possible to run in a transaction + * block if we had no XID, but this restriction is simpler for users + * to understand and we don't lose anything. */ PreventInTransactionBlock(isTopLevel, "REPACK (CONCURRENTLY)"); @@ -1001,10 +1015,8 @@ rebuild_relation(Relation OldHeap, Relation index, bool verbose, * Note that the worker has to wait for all transactions with XID * already assigned to finish. If some of those transactions is * waiting for a lock conflicting with ShareUpdateExclusiveLock on our - * table (e.g. it runs CREATE INDEX), we can end up in a deadlock. - * Not sure this risk is worth unlocking/locking the table (and its - * clustering index) and checking again if it's still eligible for - * REPACK CONCURRENTLY. + * table (e.g. it runs CREATE INDEX), it should encounter ERROR in the + * deadlock checking code. */ start_repack_decoding_worker(tableOid); @@ -3093,7 +3105,19 @@ rebuild_relation_finish_concurrent(Relation NewHeap, Relation OldHeap, LockRelationOid(OldHeap->rd_rel->reltoastrelid, AccessExclusiveLock); /* - * Tuples and pages of the old heap will be gone, but the heap will stay. + * Now that we have all access-exclusive locks on all relations, we no + * longer want other processes to error out when trying to acquire a + * conflicting lock. Therefore, unset our flag. + */ + LWLockAcquire(ProcArrayLock, LW_EXCLUSIVE); + MyProc->statusFlags &= ~PROC_IN_CONCURRENT_REPACK; + ProcGlobal->statusFlags[MyProc->pgxactoff] = MyProc->statusFlags; + LWLockRelease(ProcArrayLock); + + /* + * Tuples and pages of the old heap will be gone, but the heap itself will + * stay. In order for predicate locks to continue to work, convert them + * to relation-level locks. We do this both for table and indexes. */ TransferPredicateLocksToHeapRelation(OldHeap); foreach_ptr(RelationData, index, indexrels) @@ -3366,9 +3390,11 @@ start_repack_decoding_worker(Oid relid) /* * The decoding setup must be done before the caller can have XID assigned - * for any reason, otherwise the worker might end up in a deadlock, - * waiting for the caller's transaction to end. Therefore wait here until - * the worker indicates that it has the logical decoding initialized. + * for any reason, otherwise the worker might end up waiting for the + * caller's transaction to end. (Deadlock detector does not consider this + * a conflict because the worker is in the same locking group as the + * backend that launched it.) Therefore wait here until the worker + * indicates that it has the logical decoding initialized. */ ConditionVariablePrepareToSleep(&shared->cv); for (;;) diff --git a/src/backend/storage/lmgr/deadlock.c b/src/backend/storage/lmgr/deadlock.c index b8962d875b6..c20ac682b0d 100644 --- a/src/backend/storage/lmgr/deadlock.c +++ b/src/backend/storage/lmgr/deadlock.c @@ -620,6 +620,21 @@ FindLockCycleRecurseMember(PGPROC *checkProc, proc->statusFlags & PROC_IS_AUTOVACUUM) blocking_autovacuum_proc = proc; + /* + * Similarly, if we note that we're blocked by some + * process running REPACK (CONCURRENTLY), just fail. That + * process is going to upgrade its lock at some point, and + * it would be inappropriate for any other process to + * cause that to fail. + */ + if (checkProc == MyProc && + proc->statusFlags & PROC_IN_CONCURRENT_REPACK) + ereport(ERROR, + errcode(ERRCODE_OBJECT_IN_USE), + errmsg("could not wait for concurrent REPACK"), + errdetail("Process %d waits for REPACK running on process %d", + MyProc->pid, proc->pid)); + /* We're done looking at this proclock */ break; } diff --git a/src/include/storage/proc.h b/src/include/storage/proc.h index 1dad125706e..8ad9718f3d6 100644 --- a/src/include/storage/proc.h +++ b/src/include/storage/proc.h @@ -69,10 +69,12 @@ struct XidCache #define PROC_AFFECTS_ALL_HORIZONS 0x20 /* this proc's xmin must be * included in vacuum horizons * in all databases */ +#define PROC_IN_CONCURRENT_REPACK 0x40 /* REPACK (CONCURRENTLY) */ -/* flags reset at EOXact */ +/* flags reset at EOXact. A bit of a misnomer ... */ #define PROC_VACUUM_STATE_MASK \ - (PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND) + (PROC_IN_VACUUM | PROC_IN_SAFE_IC | PROC_VACUUM_FOR_WRAPAROUND | \ + PROC_IN_CONCURRENT_REPACK) /* * Xmin-related flags. Make sure any flags that affect how the process' Xmin diff --git a/src/test/modules/injection_points/Makefile b/src/test/modules/injection_points/Makefile index 2cd7d87c533..f7663859fe2 100644 --- a/src/test/modules/injection_points/Makefile +++ b/src/test/modules/injection_points/Makefile @@ -15,6 +15,7 @@ REGRESS_OPTS = --dlpath=$(top_builddir)/src/test/regress ISOLATION = basic \ inplace \ repack \ + repack_deadlock \ repack_toast \ syscache-update-pruned \ heap_lock_update diff --git a/src/test/modules/injection_points/expected/repack_deadlock.out b/src/test/modules/injection_points/expected/repack_deadlock.out new file mode 100644 index 00000000000..a86e4767536 --- /dev/null +++ b/src/test/modules/injection_points/expected/repack_deadlock.out @@ -0,0 +1,63 @@ +Parsed test spec with 2 sessions + +starting permutation: wait_before_lock add_column wakeup_before_lock check1 +injection_points_attach +----------------------- + +(1 row) + +step wait_before_lock: + REPACK (CONCURRENTLY) repack_deadlock USING INDEX repack_deadlock_pkey; + <waiting ...> +step add_column: + alter table repack_deadlock add column noise text; + <waiting ...> +step add_column: <... completed> +ERROR: could not wait for concurrent REPACK +step wakeup_before_lock: + SELECT injection_points_wakeup('repack-concurrently-before-lock'); + +injection_points_wakeup +----------------------- + +(1 row) + +step wait_before_lock: <... completed> +step check1: + INSERT INTO relfilenodes(node) + SELECT relfilenode FROM pg_class WHERE relname='repack_deadlock'; + + SELECT count(DISTINCT node) FROM relfilenodes; + + SELECT i, j FROM repack_deadlock ORDER BY i, j; + + INSERT INTO data_s1(i, j) + SELECT i, j FROM repack_deadlock; + + SELECT count(*) + FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j) + WHERE d1.i ISNULL OR d2.i ISNULL; + +count +----- + 1 +(1 row) + +i|j +-+- +1|1 +2|2 +3|3 +4|4 +(4 rows) + +count +----- + 4 +(1 row) + +injection_points_detach +----------------------- + +(1 row) + diff --git a/src/test/modules/injection_points/meson.build b/src/test/modules/injection_points/meson.build index a414abb924b..1cd88d6db65 100644 --- a/src/test/modules/injection_points/meson.build +++ b/src/test/modules/injection_points/meson.build @@ -46,6 +46,7 @@ tests += { 'basic', 'inplace', 'repack', + 'repack_deadlock', 'repack_toast', 'syscache-update-pruned', 'heap_lock_update', diff --git a/src/test/modules/injection_points/specs/repack_deadlock.spec b/src/test/modules/injection_points/specs/repack_deadlock.spec new file mode 100644 index 00000000000..9d23a6588c2 --- /dev/null +++ b/src/test/modules/injection_points/specs/repack_deadlock.spec @@ -0,0 +1,83 @@ +# Test REPACK with a concurrent transaction that would cause a deadlock +setup +{ + CREATE EXTENSION injection_points; + + CREATE TABLE repack_deadlock(i int PRIMARY KEY, j int); + INSERT INTO repack_deadlock(i, j) VALUES (1, 1), (2, 2), (3, 3), (4, 4); + + CREATE TABLE relfilenodes(node oid); + + CREATE TABLE data_s1(i int, j int); + CREATE TABLE data_s2(i int, j int); +} + +teardown +{ + DROP TABLE repack_deadlock; + DROP EXTENSION injection_points; + + DROP TABLE relfilenodes; + DROP TABLE data_s1; + DROP TABLE data_s2; +} + +session s1 +setup +{ + SELECT injection_points_set_local(); + SELECT injection_points_attach('repack-concurrently-before-lock', 'wait'); +} +# Perform the initial load and wait for s2 to do some data changes. +step wait_before_lock +{ + REPACK (CONCURRENTLY) repack_deadlock USING INDEX repack_deadlock_pkey; +} +# Check the table from the perspective of s1. +# +# Besides the contents, we also check that relfilenode has changed. + +# Have each session write the contents into a table and use FULL JOIN to check +# if the outputs are identical. +step check1 +{ + INSERT INTO relfilenodes(node) + SELECT relfilenode FROM pg_class WHERE relname='repack_deadlock'; + + SELECT count(DISTINCT node) FROM relfilenodes; + + SELECT i, j FROM repack_deadlock ORDER BY i, j; + + INSERT INTO data_s1(i, j) + SELECT i, j FROM repack_deadlock; + + SELECT count(*) + FROM data_s1 d1 FULL JOIN data_s2 d2 USING (i, j) + WHERE d1.i ISNULL OR d2.i ISNULL; +} +teardown +{ + SELECT injection_points_detach('repack-concurrently-before-lock'); +} + +session s2 +# Change the existing data. UPDATE changes both key and non-key columns. Also +# update one row twice to test whether tuple version generated by this session +# can be found. +step add_column +{ + alter table repack_deadlock add column noise text; +} + +step wakeup_before_lock +{ + SELECT injection_points_wakeup('repack-concurrently-before-lock'); +} + +# Test if data changes introduced while one session is performing REPACK +# CONCURRENTLY find their way into the table. +permutation + wait_before_lock + add_column + wakeup_before_lock + check1 -- 2.47.3 --gp2pyozrd5pweboh Content-Type: text/x-diff; charset=utf-8 Content-Disposition: attachment; filename="v52-0007-Check-for-transaction-block-early-in-ExecRepack.patch" ^ permalink raw reply [nested|flat] 265+ messages in thread
end of thread, other threads:[~2026-04-01 15:35 UTC | newest] Thread overview: 265+ messages (download: mbox mbox.gz follow: Atom feed) -- links below jump to the message on this page -- 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2022-01-10 19:20 [PATCH v6] README for 64bit xid Pavel Borisov <pashkin.elfe@gmail.com> 2026-04-01 15:35 [PATCH v52 06/10] Error out any process that would block at REPACK Antonin Houska <ah@cybertec.at>
This inbox is served by agora; see mirroring instructions for how to clone and mirror all data and code used for this inbox