agora inbox for pgsql-hackers@postgresql.org
help / color / mirror / Atom feedFile locks for data directory lockfile in the context of Linux namespaces
10+ messages / 3 participants
[nested] [flat]
* File locks for data directory lockfile in the context of Linux namespaces
@ 2025-12-19 14:27 Dmitry Dolgov <9erthalion6@gmail.com>
2026-01-17 15:26 ` Re: File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 1 reply; 10+ messages in thread
From: Dmitry Dolgov @ 2025-12-19 14:27 UTC (permalink / raw)
To: pgsql-hackers
Hi,
TL;DR This is a proposal to use file locking with a data directory lockfile at
startup, which helps to avoid potential Linux PID namespace visibility issues.
Recently I've stumbled upon a quite annoying problem, which will require a bit
of explanation. Currently at startup if the data directory lock exists,
postgres inspects it and assumes that if it contains the same PID as the
current process, it must be a state file after a system reboot and assigning
of the same PID again. But it seems there is another possible scenario: two
postgres instances are running concurrently inside different PID namespaces,
they don't see each other and have the same PID assigned withing the respective
namespace.
It's relatively easy to use pid/ipc/net namespaces to construct a situation,
when two postgres instances run in parallel on the same data directory and do
not notice a thing, something like this:
sudo unshare --ipc --net --pid --fork --mount-proc \
bash -c 'sudo -u postgres postgres -D data'
This of course can lead to all sorts of nasty issues, but looks very artificial
at first -- obviously whoever is responsible for namespace management must also
take care about data access isolation.
But it turns out situations like this indeed could happen in practice, when it
comes to container orchestration, mostly due to lack of knowledge or
misunderstanding of documentation. Kubernetes has one particular access mode
for volumes, ReadWriteOnce [1], which often assumed to be good enough -- but it
guarantees only a single mount per node, not per pod. Kubernetes also allows a
forced pod termination [2], which removes the pod from the API, but still gives
some grace period for the pod to finish. All of this can lead to an unfortunate
sequence of events:
* A postgres pod got forcefully terminated and removed from the API right away.
* A new postgres pod is started instead (there is nothing in the API, so
why not), while the old one is still terminating.
* If they were utilizing RWO mode, the new pod will immediately get the data
volume and can access it while the old pod is still terminating.
In the end we get a situation similar to what I've described above, and
strangely enough it looks like this indeed happens in the field.
It's fair to say that it's a Kubernetes issue (there are warnings about
that in the documentation), and PostgreSQL doesn't have anything to do
with that. But taking into account general possibility of confusing
PostgreSQL with Linux namespaces it looks to me like one of those "shoot
yourself in the foot" situation, and I became curious if there are any
easy way to improve things.
The root of the problem is lack of any time related information that PostgreSQL
could use to distinguish between two scenarios: when a single container was
killed and started again later; and when two containers run at the same time.
After some experimenting it looks like the only plausible answer could be file
locking for data directory lockfile.
This approach was discussed many times in hackers, and from what I see there
are few arguments against using file locking as the main mechanism for
protecting the data directory content:
* Portability. There are two types of file locks, advisory record locks (POSIX)
and open file description locks (was non-POSIX). The former has set of flaws,
but most importantly for this discussion is that advisory record locks are
associated with a process and thus affected by PID namespace isolation. The
later are associated with open file descriptors and are suitable solution to
fix the problem. Originally open file description locks were non-POSIX, but
looks like they have become a part of POSIX.1 2024, (see F_OFD_SETLK) [3].
* Issues with NFS. It turns out NFSv3 does not support open file description
locks and convert them into advisory locks. For our purposes it means that
the aproach will not change anything for NFSv3. Regarding NFSv4, it uses some
sort of lease system for locking, and I haven't found anything claiming that
locks will be converted to advisory.
With this in mind, it seems to me that adding file locking to data directory
lockfile as a "best efforts" approach (i.e. if it doesn't work, we continue as
before) on top of already existing mechanism will improve most of things, while
keeping status quo for some others. I've attached a quick sketch of how the
patch might look like.
Any thoughts / commentaries on that?
[1]: https://kubernetes.io/docs/concepts/storage/persistent-volumes/#access-modes
[2]: https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/#pod-termination-forced
[3]: https://pubs.opengroup.org/onlinepubs/9799919799/functions/fcntl.html
From ee18ce65f9e89a2319fd137a4ac65deb29064a84 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Thu, 18 Dec 2025 18:21:59 +0100
Subject: [PATCH v1] Use open file description locks for data directory
lockfile
When starting up, postmaster checks for an existing data directory lockfile. If
this file contains current process PID, it's assumed to be stale. Turns out
there is another possibility: we might be running in a PID namespace, and there
is another postgres running inside another PID namespace using the same data
directory. The result is that we don't see another process due to namespace
isolation and start concurrently with the other.
To prevent such situations, at startup use fcntl to get an exclusive open file
description lock for data directory lockfile. Since such locks are associated
with open file descriptors, meaning they're not affected by PID namespace
isolation. It's a "best effort" locking, intended to work with already existing
mechanism, not replace it.
This approach was discussed multiple times in the past, and usually was
rejected as the main work horse for the data directory lockfile due to:
* Portability issues. Open file description lock was a non-POSIX extension in
Linux and similar flock is from BSD standard. But looks like everybody agrees
that such locks make more sense than a typical advisory locks, and
F_OFD_SETLK made its way into POSIX.1 2024 [1].
* Issues with NFS. The current state of things here looks like this:
- NFSv3 doesn't implement open file description locks, they're converted to
advisory locks instead. Advisory locks are subject to namespace isolation,
meaning that processes in different PID namespaces will not see each other
advisory lock, and it's still possible to run multiple postgres
instances on the same data directory.
- NFSv4 uses a lease system for locking, I haven't found any mention of
conversion to advisory locks neither in the man page nor in RFC [2].
To summarize, the approach is now considered POSIX and should fix the described
problem everywhere, except NFSv3.
[1]: https://pubs.opengroup.org/onlinepubs/9799919799/functions/fcntl.html
[2]: https://www.rfc-editor.org/rfc/rfc7530
---
configure | 14 ++++
configure.ac | 3 +
meson.build | 1 +
src/backend/utils/init/miscinit.c | 107 +++++++++++++++++++++++++-----
src/include/pg_config.h.in | 4 ++
5 files changed, 111 insertions(+), 18 deletions(-)
diff --git a/configure b/configure
index 14ad0a5006f..b176ac39799 100755
--- a/configure
+++ b/configure
@@ -16177,6 +16177,20 @@ cat >>confdefs.h <<_ACEOF
_ACEOF
+# Linux open file descriptor locks
+ac_fn_c_check_decl "$LINENO" "F_OFD_SETLK" "ac_cv_have_decl_F_OFD_SETLK" "#include <fcntl.h>
+"
+if test "x$ac_cv_have_decl_F_OFD_SETLK" = xyes; then :
+ ac_have_decl=1
+else
+ ac_have_decl=0
+fi
+
+cat >>confdefs.h <<_ACEOF
+#define HAVE_DECL_F_OFD_SETLK $ac_have_decl
+_ACEOF
+
+
ac_fn_c_check_func "$LINENO" "explicit_bzero" "ac_cv_func_explicit_bzero"
if test "x$ac_cv_func_explicit_bzero" = xyes; then :
$as_echo "#define HAVE_EXPLICIT_BZERO 1" >>confdefs.h
diff --git a/configure.ac b/configure.ac
index 01b3bbc1be8..d6cf1f27771 100644
--- a/configure.ac
+++ b/configure.ac
@@ -1838,6 +1838,9 @@ AC_CHECK_DECLS([memset_s], [], [], [#define __STDC_WANT_LIB_EXT1__ 1
# This is probably only present on macOS, but may as well check always
AC_CHECK_DECLS(F_FULLFSYNC, [], [], [#include <fcntl.h>])
+# Linux open file descriptor locks
+AC_CHECK_DECLS([F_OFD_SETLK], [], [], [#include <fcntl.h>])
+
AC_REPLACE_FUNCS(m4_normalize([
explicit_bzero
getopt
diff --git a/meson.build b/meson.build
index d7c5193d4ce..ad9ecad829f 100644
--- a/meson.build
+++ b/meson.build
@@ -2667,6 +2667,7 @@ decl_checks = [
['strnlen', 'string.h'],
['strsep', 'string.h'],
['timingsafe_bcmp', 'string.h'],
+ ['F_OFD_SETLK', 'fcntl.h'],
]
# Need to check for function declarations for these functions, because
diff --git a/src/backend/utils/init/miscinit.c b/src/backend/utils/init/miscinit.c
index fec79992c8d..78bb7df543e 100644
--- a/src/backend/utils/init/miscinit.c
+++ b/src/backend/utils/init/miscinit.c
@@ -68,6 +68,11 @@ static List *lock_files = NIL;
static Latch LocalLatchData;
+#ifdef HAVE_DECL_F_OFD_SETLK
+/* File descriptor for data directory lock file. */
+static int DataDirLockFD;
+#endif
+
/* ----------------------------------------------------------------
* ignoring system indexes support stuff
*
@@ -1117,6 +1122,45 @@ RestoreClientConnectionInfo(char *conninfo)
*-------------------------------------------------------------------------
*/
+/*
+ * Flock the data directory lockfile.
+ *
+ * Lock the data directory lockfile with an open file description lock. If the
+ * lock is already taken, it's a hard stop. It's only a best effort test, and
+ * any other errors are ignored. On succes the file descriptor is duplicated,
+ * to make sure there will be at least one open copy of it to keep the lock.
+ *
+ * filename is used only for reporting purposes.
+ */
+static void
+FlockDataDirLockFile(int fd, const char *filename)
+{
+
+#ifdef HAVE_DECL_F_OFD_SETLK
+ struct flock lock;
+
+ lock.l_type = F_WRLCK;
+ lock.l_whence = SEEK_SET;
+ lock.l_start = 0;
+ lock.l_len = 0;
+ lock.l_pid = 0;
+
+ if (fcntl(fd, F_OFD_SETLK, &lock) == -1)
+ {
+ if (errno == EAGAIN)
+ ereport(FATAL,
+ (errcode(ERRCODE_LOCK_FILE_EXISTS),
+ errmsg("cannot lock the lock file \"%s\"", filename),
+ errhint("Another server is starting.")));
+ else
+ elog(WARNING, "Failed locking file \"%s\", %m", filename);
+ }
+ else
+ DataDirLockFD = dup(fd);
+#endif
+
+}
+
/*
* proc_exit callback to remove lockfiles.
*/
@@ -1125,6 +1169,11 @@ UnlinkLockFiles(int status, Datum arg)
{
ListCell *l;
+#ifdef HAVE_DECL_F_OFD_SETLK
+ /* Close the file descriptor, which keeps the open file description lock */
+ close(DataDirLockFD);
+#endif
+
foreach(l, lock_files)
{
char *curfile = (char *) lfirst(l);
@@ -1171,22 +1220,32 @@ CreateLockFile(const char *filename, bool amPostmaster,
const char *envvar;
/*
- * If the PID in the lockfile is our own PID or our parent's or
- * grandparent's PID, then the file must be stale (probably left over from
- * a previous system boot cycle). We need to check this because of the
- * likelihood that a reboot will assign exactly the same PID as we had in
- * the previous reboot, or one that's only one or two counts larger and
- * hence the lockfile's PID now refers to an ancestor shell process. We
- * allow pg_ctl to pass down its parent shell PID (our grandparent PID)
- * via the environment variable PG_GRANDPARENT_PID; this is so that
- * launching the postmaster via pg_ctl can be just as reliable as
- * launching it directly. There is no provision for detecting
- * further-removed ancestor processes, but if the init script is written
- * carefully then all but the immediate parent shell will be root-owned
- * processes and so the kill test will fail with EPERM. Note that we
- * cannot get a false negative this way, because an existing postmaster
- * would surely never launch a competing postmaster or pg_ctl process
- * directly.
+ * If we find an already existing lockfile containing our own PID,
+ * there are few options:
+ *
+ * - There is another process, that we don't see due to PID namespace
+ * isolation, which is already running in this data directory.
+ *
+ * To prevent two concurrent processes working with the same data
+ * directory, we first try to lock the lockfile exclusively.
+ *
+ * - The file must be stale, probably left over from a previous system boot
+ * cycle. The same if the lockfile contains our parent's or grandparent's
+ * PID.
+ *
+ * We need to check this because of the likelihood that a reboot will
+ * assign exactly the same PID as we had in the previous reboot, or one
+ * that's only one or two counts larger and hence the lockfile's PID now
+ * refers to an ancestor shell process. We allow pg_ctl to pass down its
+ * parent shell PID (our grandparent PID) via the environment variable
+ * PG_GRANDPARENT_PID; this is so that launching the postmaster via
+ * pg_ctl can be just as reliable as launching it directly. There is no
+ * provision for detecting further-removed ancestor processes, but if the
+ * init script is written carefully then all but the immediate parent
+ * shell will be root-owned processes and so the kill test will fail with
+ * EPERM. Note that we cannot get a false negative this way, because an
+ * existing postmaster would surely never launch a competing postmaster
+ * or pg_ctl process directly.
*/
my_pid = getpid();
@@ -1222,7 +1281,11 @@ CreateLockFile(const char *filename, bool amPostmaster,
*/
fd = open(filename, O_RDWR | O_CREAT | O_EXCL, pg_file_create_mode);
if (fd >= 0)
- break; /* Success; exit the retry loop */
+ {
+ /* Success; lock and exit the retry loop */
+ FlockDataDirLockFile(fd, filename);
+ break;
+ }
/*
* Couldn't create the pid file. Probably it already exists.
@@ -1236,8 +1299,12 @@ CreateLockFile(const char *filename, bool amPostmaster,
/*
* Read the file to get the old owner's PID. Note race condition
* here: file might have been deleted since we tried to create it.
+ *
+ * We're going to use the same fd for flock, and want to create a write
+ * lock for the latter one. Since both fd and the lock have to be of
+ * the same type, open the file for read and write.
*/
- fd = open(filename, O_RDONLY, pg_file_create_mode);
+ fd = open(filename, O_RDWR, pg_file_create_mode);
if (fd < 0)
{
if (errno == ENOENT)
@@ -1247,6 +1314,10 @@ CreateLockFile(const char *filename, bool amPostmaster,
errmsg("could not open lock file \"%s\": %m",
filename)));
}
+
+ /* Try to lock the file. We stop here, if it's already locked. */
+ FlockDataDirLockFile(fd, filename);
+
pgstat_report_wait_start(WAIT_EVENT_LOCK_FILE_CREATE_READ);
if ((len = read(fd, buffer, sizeof(buffer) - 1)) < 0)
ereport(FATAL,
diff --git a/src/include/pg_config.h.in b/src/include/pg_config.h.in
index 92fcc5f3063..c19a50f108e 100644
--- a/src/include/pg_config.h.in
+++ b/src/include/pg_config.h.in
@@ -83,6 +83,10 @@
don't. */
#undef HAVE_DECL_F_FULLFSYNC
+/* Define to 1 if you have the declaration of `F_OFD_SETLK', and to 0 if you
+ don't. */
+#undef HAVE_DECL_F_OFD_SETLK
+
/* Define to 1 if you have the declaration of
`LLVMCreateGDBRegistrationListener', and to 0 if you don't. */
#undef HAVE_DECL_LLVMCREATEGDBREGISTRATIONLISTENER
base-commit: 44f49511b7940adf3be4337d4feb2de38fe92297
--
2.49.0
Attachments:
[text/plain] v1-0001-Use-open-file-description-locks-for-data-director.patch (10.6K, ../../z2u3jtjzhqofqhrjvfgkwl4cczhpqibz3t47gpx7fbe7ceaihk@i5mveu2wfpxn/2-v1-0001-Use-open-file-description-locks-for-data-director.patch)
download | inline diff:
From ee18ce65f9e89a2319fd137a4ac65deb29064a84 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Thu, 18 Dec 2025 18:21:59 +0100
Subject: [PATCH v1] Use open file description locks for data directory
lockfile
When starting up, postmaster checks for an existing data directory lockfile. If
this file contains current process PID, it's assumed to be stale. Turns out
there is another possibility: we might be running in a PID namespace, and there
is another postgres running inside another PID namespace using the same data
directory. The result is that we don't see another process due to namespace
isolation and start concurrently with the other.
To prevent such situations, at startup use fcntl to get an exclusive open file
description lock for data directory lockfile. Since such locks are associated
with open file descriptors, meaning they're not affected by PID namespace
isolation. It's a "best effort" locking, intended to work with already existing
mechanism, not replace it.
This approach was discussed multiple times in the past, and usually was
rejected as the main work horse for the data directory lockfile due to:
* Portability issues. Open file description lock was a non-POSIX extension in
Linux and similar flock is from BSD standard. But looks like everybody agrees
that such locks make more sense than a typical advisory locks, and
F_OFD_SETLK made its way into POSIX.1 2024 [1].
* Issues with NFS. The current state of things here looks like this:
- NFSv3 doesn't implement open file description locks, they're converted to
advisory locks instead. Advisory locks are subject to namespace isolation,
meaning that processes in different PID namespaces will not see each other
advisory lock, and it's still possible to run multiple postgres
instances on the same data directory.
- NFSv4 uses a lease system for locking, I haven't found any mention of
conversion to advisory locks neither in the man page nor in RFC [2].
To summarize, the approach is now considered POSIX and should fix the described
problem everywhere, except NFSv3.
[1]: https://pubs.opengroup.org/onlinepubs/9799919799/functions/fcntl.html
[2]: https://www.rfc-editor.org/rfc/rfc7530
---
configure | 14 ++++
configure.ac | 3 +
meson.build | 1 +
src/backend/utils/init/miscinit.c | 107 +++++++++++++++++++++++++-----
src/include/pg_config.h.in | 4 ++
5 files changed, 111 insertions(+), 18 deletions(-)
diff --git a/configure b/configure
index 14ad0a5006f..b176ac39799 100755
--- a/configure
+++ b/configure
@@ -16177,6 +16177,20 @@ cat >>confdefs.h <<_ACEOF
_ACEOF
+# Linux open file descriptor locks
+ac_fn_c_check_decl "$LINENO" "F_OFD_SETLK" "ac_cv_have_decl_F_OFD_SETLK" "#include <fcntl.h>
+"
+if test "x$ac_cv_have_decl_F_OFD_SETLK" = xyes; then :
+ ac_have_decl=1
+else
+ ac_have_decl=0
+fi
+
+cat >>confdefs.h <<_ACEOF
+#define HAVE_DECL_F_OFD_SETLK $ac_have_decl
+_ACEOF
+
+
ac_fn_c_check_func "$LINENO" "explicit_bzero" "ac_cv_func_explicit_bzero"
if test "x$ac_cv_func_explicit_bzero" = xyes; then :
$as_echo "#define HAVE_EXPLICIT_BZERO 1" >>confdefs.h
diff --git a/configure.ac b/configure.ac
index 01b3bbc1be8..d6cf1f27771 100644
--- a/configure.ac
+++ b/configure.ac
@@ -1838,6 +1838,9 @@ AC_CHECK_DECLS([memset_s], [], [], [#define __STDC_WANT_LIB_EXT1__ 1
# This is probably only present on macOS, but may as well check always
AC_CHECK_DECLS(F_FULLFSYNC, [], [], [#include <fcntl.h>])
+# Linux open file descriptor locks
+AC_CHECK_DECLS([F_OFD_SETLK], [], [], [#include <fcntl.h>])
+
AC_REPLACE_FUNCS(m4_normalize([
explicit_bzero
getopt
diff --git a/meson.build b/meson.build
index d7c5193d4ce..ad9ecad829f 100644
--- a/meson.build
+++ b/meson.build
@@ -2667,6 +2667,7 @@ decl_checks = [
['strnlen', 'string.h'],
['strsep', 'string.h'],
['timingsafe_bcmp', 'string.h'],
+ ['F_OFD_SETLK', 'fcntl.h'],
]
# Need to check for function declarations for these functions, because
diff --git a/src/backend/utils/init/miscinit.c b/src/backend/utils/init/miscinit.c
index fec79992c8d..78bb7df543e 100644
--- a/src/backend/utils/init/miscinit.c
+++ b/src/backend/utils/init/miscinit.c
@@ -68,6 +68,11 @@ static List *lock_files = NIL;
static Latch LocalLatchData;
+#ifdef HAVE_DECL_F_OFD_SETLK
+/* File descriptor for data directory lock file. */
+static int DataDirLockFD;
+#endif
+
/* ----------------------------------------------------------------
* ignoring system indexes support stuff
*
@@ -1117,6 +1122,45 @@ RestoreClientConnectionInfo(char *conninfo)
*-------------------------------------------------------------------------
*/
+/*
+ * Flock the data directory lockfile.
+ *
+ * Lock the data directory lockfile with an open file description lock. If the
+ * lock is already taken, it's a hard stop. It's only a best effort test, and
+ * any other errors are ignored. On succes the file descriptor is duplicated,
+ * to make sure there will be at least one open copy of it to keep the lock.
+ *
+ * filename is used only for reporting purposes.
+ */
+static void
+FlockDataDirLockFile(int fd, const char *filename)
+{
+
+#ifdef HAVE_DECL_F_OFD_SETLK
+ struct flock lock;
+
+ lock.l_type = F_WRLCK;
+ lock.l_whence = SEEK_SET;
+ lock.l_start = 0;
+ lock.l_len = 0;
+ lock.l_pid = 0;
+
+ if (fcntl(fd, F_OFD_SETLK, &lock) == -1)
+ {
+ if (errno == EAGAIN)
+ ereport(FATAL,
+ (errcode(ERRCODE_LOCK_FILE_EXISTS),
+ errmsg("cannot lock the lock file \"%s\"", filename),
+ errhint("Another server is starting.")));
+ else
+ elog(WARNING, "Failed locking file \"%s\", %m", filename);
+ }
+ else
+ DataDirLockFD = dup(fd);
+#endif
+
+}
+
/*
* proc_exit callback to remove lockfiles.
*/
@@ -1125,6 +1169,11 @@ UnlinkLockFiles(int status, Datum arg)
{
ListCell *l;
+#ifdef HAVE_DECL_F_OFD_SETLK
+ /* Close the file descriptor, which keeps the open file description lock */
+ close(DataDirLockFD);
+#endif
+
foreach(l, lock_files)
{
char *curfile = (char *) lfirst(l);
@@ -1171,22 +1220,32 @@ CreateLockFile(const char *filename, bool amPostmaster,
const char *envvar;
/*
- * If the PID in the lockfile is our own PID or our parent's or
- * grandparent's PID, then the file must be stale (probably left over from
- * a previous system boot cycle). We need to check this because of the
- * likelihood that a reboot will assign exactly the same PID as we had in
- * the previous reboot, or one that's only one or two counts larger and
- * hence the lockfile's PID now refers to an ancestor shell process. We
- * allow pg_ctl to pass down its parent shell PID (our grandparent PID)
- * via the environment variable PG_GRANDPARENT_PID; this is so that
- * launching the postmaster via pg_ctl can be just as reliable as
- * launching it directly. There is no provision for detecting
- * further-removed ancestor processes, but if the init script is written
- * carefully then all but the immediate parent shell will be root-owned
- * processes and so the kill test will fail with EPERM. Note that we
- * cannot get a false negative this way, because an existing postmaster
- * would surely never launch a competing postmaster or pg_ctl process
- * directly.
+ * If we find an already existing lockfile containing our own PID,
+ * there are few options:
+ *
+ * - There is another process, that we don't see due to PID namespace
+ * isolation, which is already running in this data directory.
+ *
+ * To prevent two concurrent processes working with the same data
+ * directory, we first try to lock the lockfile exclusively.
+ *
+ * - The file must be stale, probably left over from a previous system boot
+ * cycle. The same if the lockfile contains our parent's or grandparent's
+ * PID.
+ *
+ * We need to check this because of the likelihood that a reboot will
+ * assign exactly the same PID as we had in the previous reboot, or one
+ * that's only one or two counts larger and hence the lockfile's PID now
+ * refers to an ancestor shell process. We allow pg_ctl to pass down its
+ * parent shell PID (our grandparent PID) via the environment variable
+ * PG_GRANDPARENT_PID; this is so that launching the postmaster via
+ * pg_ctl can be just as reliable as launching it directly. There is no
+ * provision for detecting further-removed ancestor processes, but if the
+ * init script is written carefully then all but the immediate parent
+ * shell will be root-owned processes and so the kill test will fail with
+ * EPERM. Note that we cannot get a false negative this way, because an
+ * existing postmaster would surely never launch a competing postmaster
+ * or pg_ctl process directly.
*/
my_pid = getpid();
@@ -1222,7 +1281,11 @@ CreateLockFile(const char *filename, bool amPostmaster,
*/
fd = open(filename, O_RDWR | O_CREAT | O_EXCL, pg_file_create_mode);
if (fd >= 0)
- break; /* Success; exit the retry loop */
+ {
+ /* Success; lock and exit the retry loop */
+ FlockDataDirLockFile(fd, filename);
+ break;
+ }
/*
* Couldn't create the pid file. Probably it already exists.
@@ -1236,8 +1299,12 @@ CreateLockFile(const char *filename, bool amPostmaster,
/*
* Read the file to get the old owner's PID. Note race condition
* here: file might have been deleted since we tried to create it.
+ *
+ * We're going to use the same fd for flock, and want to create a write
+ * lock for the latter one. Since both fd and the lock have to be of
+ * the same type, open the file for read and write.
*/
- fd = open(filename, O_RDONLY, pg_file_create_mode);
+ fd = open(filename, O_RDWR, pg_file_create_mode);
if (fd < 0)
{
if (errno == ENOENT)
@@ -1247,6 +1314,10 @@ CreateLockFile(const char *filename, bool amPostmaster,
errmsg("could not open lock file \"%s\": %m",
filename)));
}
+
+ /* Try to lock the file. We stop here, if it's already locked. */
+ FlockDataDirLockFile(fd, filename);
+
pgstat_report_wait_start(WAIT_EVENT_LOCK_FILE_CREATE_READ);
if ((len = read(fd, buffer, sizeof(buffer) - 1)) < 0)
ereport(FATAL,
diff --git a/src/include/pg_config.h.in b/src/include/pg_config.h.in
index 92fcc5f3063..c19a50f108e 100644
--- a/src/include/pg_config.h.in
+++ b/src/include/pg_config.h.in
@@ -83,6 +83,10 @@
don't. */
#undef HAVE_DECL_F_FULLFSYNC
+/* Define to 1 if you have the declaration of `F_OFD_SETLK', and to 0 if you
+ don't. */
+#undef HAVE_DECL_F_OFD_SETLK
+
/* Define to 1 if you have the declaration of
`LLVMCreateGDBRegistrationListener', and to 0 if you don't. */
#undef HAVE_DECL_LLVMCREATEGDBREGISTRATIONLISTENER
base-commit: 44f49511b7940adf3be4337d4feb2de38fe92297
--
2.49.0
^ permalink raw reply [nested|flat] 10+ messages in thread
* Re: File locks for data directory lockfile in the context of Linux namespaces
2025-12-19 14:27 File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
@ 2026-01-17 15:26 ` Dmitry Dolgov <9erthalion6@gmail.com>
2026-06-05 12:37 ` Re: File locks for data directory lockfile in the context of Linux namespaces Ilmar Yunusov <tanswis42@gmail.com>
0 siblings, 1 reply; 10+ messages in thread
From: Dmitry Dolgov @ 2026-01-17 15:26 UTC (permalink / raw)
To: pgsql-hackers
> On Fri, Dec 19, 2025 at 03:27:40PM +0100, Dmitry Dolgov wrote:
> Hi,
>
> TL;DR This is a proposal to use file locking with a data directory lockfile at
> startup, which helps to avoid potential Linux PID namespace visibility issues.
Rebased and fixed a silly error with the flag availability. Feedback is
still welcome.
From 8870df539ca25fcd7ed10f778d61fbcadf4ed5d6 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Thu, 18 Dec 2025 18:21:59 +0100
Subject: [PATCH v2] Use open file description locks for data directory
lockfile
When starting up, postmaster checks for an existing data directory lockfile. If
this file contains current process PID, it's assumed to be stale. Turns out
there is another possibility: we might be running in a PID namespace, and there
is another postgres running inside another PID namespace using the same data
directory. The result is that we don't see another process due to namespace
isolation and start concurrently with the other.
To prevent such situations, at startup use fcntl to get an exclusive open file
description lock for data directory lockfile. Since such locks are associated
with open file descriptors, meaning they're not affected by PID namespace
isolation. It's a "best effort" locking, intended to work with already existing
mechanism, not replace it.
This approach was discussed multiple times in the past, and usually was
rejected as the main work horse for the data directory lockfile due to:
* Portability issues. Open file description lock was a non-POSIX extension in
Linux and similar flock is from BSD standard. But looks like everybody agrees
that such locks make more sense than a typical advisory locks, and
F_OFD_SETLK made its way into POSIX.1 2024 [1].
* Issues with NFS. The current state of things here looks like this:
- NFSv3 doesn't implement open file description locks, they're converted to
advisory locks instead. Advisory locks are subject to namespace isolation,
meaning that processes in different PID namespaces will not see each other
advisory lock, and it's still possible to run multiple postgres
instances on the same data directory.
- NFSv4 uses a lease system for locking, I haven't found any mention of
conversion to advisory locks neither in the man page nor in RFC [2].
To summarize, the approach is now considered POSIX and should fix the described
problem everywhere, except NFSv3.
[1]: https://pubs.opengroup.org/onlinepubs/9799919799/functions/fcntl.html
[2]: https://www.rfc-editor.org/rfc/rfc7530
---
configure | 14 ++++
configure.ac | 3 +
meson.build | 1 +
src/backend/utils/init/miscinit.c | 107 +++++++++++++++++++++++++-----
src/include/pg_config.h.in | 4 ++
5 files changed, 111 insertions(+), 18 deletions(-)
diff --git a/configure b/configure
index fb6a4914b06..1da073f0aaa 100755
--- a/configure
+++ b/configure
@@ -16177,6 +16177,20 @@ cat >>confdefs.h <<_ACEOF
_ACEOF
+# Linux open file descriptor locks
+ac_fn_c_check_decl "$LINENO" "F_OFD_SETLK" "ac_cv_have_decl_F_OFD_SETLK" "#include <fcntl.h>
+"
+if test "x$ac_cv_have_decl_F_OFD_SETLK" = xyes; then :
+ ac_have_decl=1
+else
+ ac_have_decl=0
+fi
+
+cat >>confdefs.h <<_ACEOF
+#define HAVE_DECL_F_OFD_SETLK $ac_have_decl
+_ACEOF
+
+
ac_fn_c_check_func "$LINENO" "explicit_bzero" "ac_cv_func_explicit_bzero"
if test "x$ac_cv_func_explicit_bzero" = xyes; then :
$as_echo "#define HAVE_EXPLICIT_BZERO 1" >>confdefs.h
diff --git a/configure.ac b/configure.ac
index d3febfe58f1..075194a8af8 100644
--- a/configure.ac
+++ b/configure.ac
@@ -1844,6 +1844,9 @@ AC_CHECK_DECLS([memset_s], [], [], [#define __STDC_WANT_LIB_EXT1__ 1
# This is probably only present on macOS, but may as well check always
AC_CHECK_DECLS(F_FULLFSYNC, [], [], [#include <fcntl.h>])
+# Linux open file descriptor locks
+AC_CHECK_DECLS([F_OFD_SETLK], [], [], [#include <fcntl.h>])
+
AC_REPLACE_FUNCS(m4_normalize([
explicit_bzero
getopt
diff --git a/meson.build b/meson.build
index 6d304f32fb0..d346aefe9b6 100644
--- a/meson.build
+++ b/meson.build
@@ -2694,6 +2694,7 @@ decl_checks = [
['strlcpy', 'string.h'],
['strsep', 'string.h'],
['timingsafe_bcmp', 'string.h'],
+ ['F_OFD_SETLK', 'fcntl.h'],
]
# Need to check for function declarations for these functions, because
diff --git a/src/backend/utils/init/miscinit.c b/src/backend/utils/init/miscinit.c
index 563f20374ff..b41a5f9dcca 100644
--- a/src/backend/utils/init/miscinit.c
+++ b/src/backend/utils/init/miscinit.c
@@ -68,6 +68,11 @@ static List *lock_files = NIL;
static Latch LocalLatchData;
+#if HAVE_DECL_F_OFD_SETLK
+/* File descriptor for data directory lock file. */
+static int DataDirLockFD;
+#endif
+
/* ----------------------------------------------------------------
* ignoring system indexes support stuff
*
@@ -1117,6 +1122,45 @@ RestoreClientConnectionInfo(char *conninfo)
*-------------------------------------------------------------------------
*/
+/*
+ * Flock the data directory lockfile.
+ *
+ * Lock the data directory lockfile with an open file description lock. If the
+ * lock is already taken, it's a hard stop. It's only a best effort test, and
+ * any other errors are ignored. On succes the file descriptor is duplicated,
+ * to make sure there will be at least one open copy of it to keep the lock.
+ *
+ * filename is used only for reporting purposes.
+ */
+static void
+FlockDataDirLockFile(int fd, const char *filename)
+{
+
+#if HAVE_DECL_F_OFD_SETLK
+ struct flock lock;
+
+ lock.l_type = F_WRLCK;
+ lock.l_whence = SEEK_SET;
+ lock.l_start = 0;
+ lock.l_len = 0;
+ lock.l_pid = 0;
+
+ if (fcntl(fd, F_OFD_SETLK, &lock) == -1)
+ {
+ if (errno == EAGAIN)
+ ereport(FATAL,
+ (errcode(ERRCODE_LOCK_FILE_EXISTS),
+ errmsg("cannot lock the lock file \"%s\"", filename),
+ errhint("Another server is starting.")));
+ else
+ elog(WARNING, "Failed locking file \"%s\", %m", filename);
+ }
+ else
+ DataDirLockFD = dup(fd);
+#endif
+
+}
+
/*
* proc_exit callback to remove lockfiles.
*/
@@ -1125,6 +1169,11 @@ UnlinkLockFiles(int status, Datum arg)
{
ListCell *l;
+#if HAVE_DECL_F_OFD_SETLK
+ /* Close the file descriptor, which keeps the open file description lock */
+ close(DataDirLockFD);
+#endif
+
foreach(l, lock_files)
{
char *curfile = (char *) lfirst(l);
@@ -1171,22 +1220,32 @@ CreateLockFile(const char *filename, bool amPostmaster,
const char *envvar;
/*
- * If the PID in the lockfile is our own PID or our parent's or
- * grandparent's PID, then the file must be stale (probably left over from
- * a previous system boot cycle). We need to check this because of the
- * likelihood that a reboot will assign exactly the same PID as we had in
- * the previous reboot, or one that's only one or two counts larger and
- * hence the lockfile's PID now refers to an ancestor shell process. We
- * allow pg_ctl to pass down its parent shell PID (our grandparent PID)
- * via the environment variable PG_GRANDPARENT_PID; this is so that
- * launching the postmaster via pg_ctl can be just as reliable as
- * launching it directly. There is no provision for detecting
- * further-removed ancestor processes, but if the init script is written
- * carefully then all but the immediate parent shell will be root-owned
- * processes and so the kill test will fail with EPERM. Note that we
- * cannot get a false negative this way, because an existing postmaster
- * would surely never launch a competing postmaster or pg_ctl process
- * directly.
+ * If we find an already existing lockfile containing our own PID,
+ * there are few options:
+ *
+ * - There is another process, that we don't see due to PID namespace
+ * isolation, which is already running in this data directory.
+ *
+ * To prevent two concurrent processes working with the same data
+ * directory, we first try to lock the lockfile exclusively.
+ *
+ * - The file must be stale, probably left over from a previous system boot
+ * cycle. The same if the lockfile contains our parent's or grandparent's
+ * PID.
+ *
+ * We need to check this because of the likelihood that a reboot will
+ * assign exactly the same PID as we had in the previous reboot, or one
+ * that's only one or two counts larger and hence the lockfile's PID now
+ * refers to an ancestor shell process. We allow pg_ctl to pass down its
+ * parent shell PID (our grandparent PID) via the environment variable
+ * PG_GRANDPARENT_PID; this is so that launching the postmaster via
+ * pg_ctl can be just as reliable as launching it directly. There is no
+ * provision for detecting further-removed ancestor processes, but if the
+ * init script is written carefully then all but the immediate parent
+ * shell will be root-owned processes and so the kill test will fail with
+ * EPERM. Note that we cannot get a false negative this way, because an
+ * existing postmaster would surely never launch a competing postmaster
+ * or pg_ctl process directly.
*/
my_pid = getpid();
@@ -1222,7 +1281,11 @@ CreateLockFile(const char *filename, bool amPostmaster,
*/
fd = open(filename, O_RDWR | O_CREAT | O_EXCL, pg_file_create_mode);
if (fd >= 0)
- break; /* Success; exit the retry loop */
+ {
+ /* Success; lock and exit the retry loop */
+ FlockDataDirLockFile(fd, filename);
+ break;
+ }
/*
* Couldn't create the pid file. Probably it already exists.
@@ -1236,8 +1299,12 @@ CreateLockFile(const char *filename, bool amPostmaster,
/*
* Read the file to get the old owner's PID. Note race condition
* here: file might have been deleted since we tried to create it.
+ *
+ * We're going to use the same fd for flock, and want to create a write
+ * lock for the latter one. Since both fd and the lock have to be of
+ * the same type, open the file for read and write.
*/
- fd = open(filename, O_RDONLY, pg_file_create_mode);
+ fd = open(filename, O_RDWR, pg_file_create_mode);
if (fd < 0)
{
if (errno == ENOENT)
@@ -1247,6 +1314,10 @@ CreateLockFile(const char *filename, bool amPostmaster,
errmsg("could not open lock file \"%s\": %m",
filename)));
}
+
+ /* Try to lock the file. We stop here, if it's already locked. */
+ FlockDataDirLockFile(fd, filename);
+
pgstat_report_wait_start(WAIT_EVENT_LOCK_FILE_CREATE_READ);
if ((len = read(fd, buffer, sizeof(buffer) - 1)) < 0)
ereport(FATAL,
diff --git a/src/include/pg_config.h.in b/src/include/pg_config.h.in
index 339268dc8ef..e7b8e023829 100644
--- a/src/include/pg_config.h.in
+++ b/src/include/pg_config.h.in
@@ -80,6 +80,10 @@
don't. */
#undef HAVE_DECL_F_FULLFSYNC
+/* Define to 1 if you have the declaration of `F_OFD_SETLK', and to 0 if you
+ don't. */
+#undef HAVE_DECL_F_OFD_SETLK
+
/* Define to 1 if you have the declaration of `memset_s', and to 0 if you
don't. */
#undef HAVE_DECL_MEMSET_S
base-commit: 6831cd9e3b082d7b830c3196742dd49e3540c49b
--
2.49.0
Attachments:
[text/plain] v2-0001-Use-open-file-description-locks-for-data-director.patch (10.5K, ../../45swup77t7utfy33lu47mvkb4exzm3oeakz25abqxuvmsyzt3c@5umrpfgc3ypg/2-v2-0001-Use-open-file-description-locks-for-data-director.patch)
download | inline diff:
From 8870df539ca25fcd7ed10f778d61fbcadf4ed5d6 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Thu, 18 Dec 2025 18:21:59 +0100
Subject: [PATCH v2] Use open file description locks for data directory
lockfile
When starting up, postmaster checks for an existing data directory lockfile. If
this file contains current process PID, it's assumed to be stale. Turns out
there is another possibility: we might be running in a PID namespace, and there
is another postgres running inside another PID namespace using the same data
directory. The result is that we don't see another process due to namespace
isolation and start concurrently with the other.
To prevent such situations, at startup use fcntl to get an exclusive open file
description lock for data directory lockfile. Since such locks are associated
with open file descriptors, meaning they're not affected by PID namespace
isolation. It's a "best effort" locking, intended to work with already existing
mechanism, not replace it.
This approach was discussed multiple times in the past, and usually was
rejected as the main work horse for the data directory lockfile due to:
* Portability issues. Open file description lock was a non-POSIX extension in
Linux and similar flock is from BSD standard. But looks like everybody agrees
that such locks make more sense than a typical advisory locks, and
F_OFD_SETLK made its way into POSIX.1 2024 [1].
* Issues with NFS. The current state of things here looks like this:
- NFSv3 doesn't implement open file description locks, they're converted to
advisory locks instead. Advisory locks are subject to namespace isolation,
meaning that processes in different PID namespaces will not see each other
advisory lock, and it's still possible to run multiple postgres
instances on the same data directory.
- NFSv4 uses a lease system for locking, I haven't found any mention of
conversion to advisory locks neither in the man page nor in RFC [2].
To summarize, the approach is now considered POSIX and should fix the described
problem everywhere, except NFSv3.
[1]: https://pubs.opengroup.org/onlinepubs/9799919799/functions/fcntl.html
[2]: https://www.rfc-editor.org/rfc/rfc7530
---
configure | 14 ++++
configure.ac | 3 +
meson.build | 1 +
src/backend/utils/init/miscinit.c | 107 +++++++++++++++++++++++++-----
src/include/pg_config.h.in | 4 ++
5 files changed, 111 insertions(+), 18 deletions(-)
diff --git a/configure b/configure
index fb6a4914b06..1da073f0aaa 100755
--- a/configure
+++ b/configure
@@ -16177,6 +16177,20 @@ cat >>confdefs.h <<_ACEOF
_ACEOF
+# Linux open file descriptor locks
+ac_fn_c_check_decl "$LINENO" "F_OFD_SETLK" "ac_cv_have_decl_F_OFD_SETLK" "#include <fcntl.h>
+"
+if test "x$ac_cv_have_decl_F_OFD_SETLK" = xyes; then :
+ ac_have_decl=1
+else
+ ac_have_decl=0
+fi
+
+cat >>confdefs.h <<_ACEOF
+#define HAVE_DECL_F_OFD_SETLK $ac_have_decl
+_ACEOF
+
+
ac_fn_c_check_func "$LINENO" "explicit_bzero" "ac_cv_func_explicit_bzero"
if test "x$ac_cv_func_explicit_bzero" = xyes; then :
$as_echo "#define HAVE_EXPLICIT_BZERO 1" >>confdefs.h
diff --git a/configure.ac b/configure.ac
index d3febfe58f1..075194a8af8 100644
--- a/configure.ac
+++ b/configure.ac
@@ -1844,6 +1844,9 @@ AC_CHECK_DECLS([memset_s], [], [], [#define __STDC_WANT_LIB_EXT1__ 1
# This is probably only present on macOS, but may as well check always
AC_CHECK_DECLS(F_FULLFSYNC, [], [], [#include <fcntl.h>])
+# Linux open file descriptor locks
+AC_CHECK_DECLS([F_OFD_SETLK], [], [], [#include <fcntl.h>])
+
AC_REPLACE_FUNCS(m4_normalize([
explicit_bzero
getopt
diff --git a/meson.build b/meson.build
index 6d304f32fb0..d346aefe9b6 100644
--- a/meson.build
+++ b/meson.build
@@ -2694,6 +2694,7 @@ decl_checks = [
['strlcpy', 'string.h'],
['strsep', 'string.h'],
['timingsafe_bcmp', 'string.h'],
+ ['F_OFD_SETLK', 'fcntl.h'],
]
# Need to check for function declarations for these functions, because
diff --git a/src/backend/utils/init/miscinit.c b/src/backend/utils/init/miscinit.c
index 563f20374ff..b41a5f9dcca 100644
--- a/src/backend/utils/init/miscinit.c
+++ b/src/backend/utils/init/miscinit.c
@@ -68,6 +68,11 @@ static List *lock_files = NIL;
static Latch LocalLatchData;
+#if HAVE_DECL_F_OFD_SETLK
+/* File descriptor for data directory lock file. */
+static int DataDirLockFD;
+#endif
+
/* ----------------------------------------------------------------
* ignoring system indexes support stuff
*
@@ -1117,6 +1122,45 @@ RestoreClientConnectionInfo(char *conninfo)
*-------------------------------------------------------------------------
*/
+/*
+ * Flock the data directory lockfile.
+ *
+ * Lock the data directory lockfile with an open file description lock. If the
+ * lock is already taken, it's a hard stop. It's only a best effort test, and
+ * any other errors are ignored. On succes the file descriptor is duplicated,
+ * to make sure there will be at least one open copy of it to keep the lock.
+ *
+ * filename is used only for reporting purposes.
+ */
+static void
+FlockDataDirLockFile(int fd, const char *filename)
+{
+
+#if HAVE_DECL_F_OFD_SETLK
+ struct flock lock;
+
+ lock.l_type = F_WRLCK;
+ lock.l_whence = SEEK_SET;
+ lock.l_start = 0;
+ lock.l_len = 0;
+ lock.l_pid = 0;
+
+ if (fcntl(fd, F_OFD_SETLK, &lock) == -1)
+ {
+ if (errno == EAGAIN)
+ ereport(FATAL,
+ (errcode(ERRCODE_LOCK_FILE_EXISTS),
+ errmsg("cannot lock the lock file \"%s\"", filename),
+ errhint("Another server is starting.")));
+ else
+ elog(WARNING, "Failed locking file \"%s\", %m", filename);
+ }
+ else
+ DataDirLockFD = dup(fd);
+#endif
+
+}
+
/*
* proc_exit callback to remove lockfiles.
*/
@@ -1125,6 +1169,11 @@ UnlinkLockFiles(int status, Datum arg)
{
ListCell *l;
+#if HAVE_DECL_F_OFD_SETLK
+ /* Close the file descriptor, which keeps the open file description lock */
+ close(DataDirLockFD);
+#endif
+
foreach(l, lock_files)
{
char *curfile = (char *) lfirst(l);
@@ -1171,22 +1220,32 @@ CreateLockFile(const char *filename, bool amPostmaster,
const char *envvar;
/*
- * If the PID in the lockfile is our own PID or our parent's or
- * grandparent's PID, then the file must be stale (probably left over from
- * a previous system boot cycle). We need to check this because of the
- * likelihood that a reboot will assign exactly the same PID as we had in
- * the previous reboot, or one that's only one or two counts larger and
- * hence the lockfile's PID now refers to an ancestor shell process. We
- * allow pg_ctl to pass down its parent shell PID (our grandparent PID)
- * via the environment variable PG_GRANDPARENT_PID; this is so that
- * launching the postmaster via pg_ctl can be just as reliable as
- * launching it directly. There is no provision for detecting
- * further-removed ancestor processes, but if the init script is written
- * carefully then all but the immediate parent shell will be root-owned
- * processes and so the kill test will fail with EPERM. Note that we
- * cannot get a false negative this way, because an existing postmaster
- * would surely never launch a competing postmaster or pg_ctl process
- * directly.
+ * If we find an already existing lockfile containing our own PID,
+ * there are few options:
+ *
+ * - There is another process, that we don't see due to PID namespace
+ * isolation, which is already running in this data directory.
+ *
+ * To prevent two concurrent processes working with the same data
+ * directory, we first try to lock the lockfile exclusively.
+ *
+ * - The file must be stale, probably left over from a previous system boot
+ * cycle. The same if the lockfile contains our parent's or grandparent's
+ * PID.
+ *
+ * We need to check this because of the likelihood that a reboot will
+ * assign exactly the same PID as we had in the previous reboot, or one
+ * that's only one or two counts larger and hence the lockfile's PID now
+ * refers to an ancestor shell process. We allow pg_ctl to pass down its
+ * parent shell PID (our grandparent PID) via the environment variable
+ * PG_GRANDPARENT_PID; this is so that launching the postmaster via
+ * pg_ctl can be just as reliable as launching it directly. There is no
+ * provision for detecting further-removed ancestor processes, but if the
+ * init script is written carefully then all but the immediate parent
+ * shell will be root-owned processes and so the kill test will fail with
+ * EPERM. Note that we cannot get a false negative this way, because an
+ * existing postmaster would surely never launch a competing postmaster
+ * or pg_ctl process directly.
*/
my_pid = getpid();
@@ -1222,7 +1281,11 @@ CreateLockFile(const char *filename, bool amPostmaster,
*/
fd = open(filename, O_RDWR | O_CREAT | O_EXCL, pg_file_create_mode);
if (fd >= 0)
- break; /* Success; exit the retry loop */
+ {
+ /* Success; lock and exit the retry loop */
+ FlockDataDirLockFile(fd, filename);
+ break;
+ }
/*
* Couldn't create the pid file. Probably it already exists.
@@ -1236,8 +1299,12 @@ CreateLockFile(const char *filename, bool amPostmaster,
/*
* Read the file to get the old owner's PID. Note race condition
* here: file might have been deleted since we tried to create it.
+ *
+ * We're going to use the same fd for flock, and want to create a write
+ * lock for the latter one. Since both fd and the lock have to be of
+ * the same type, open the file for read and write.
*/
- fd = open(filename, O_RDONLY, pg_file_create_mode);
+ fd = open(filename, O_RDWR, pg_file_create_mode);
if (fd < 0)
{
if (errno == ENOENT)
@@ -1247,6 +1314,10 @@ CreateLockFile(const char *filename, bool amPostmaster,
errmsg("could not open lock file \"%s\": %m",
filename)));
}
+
+ /* Try to lock the file. We stop here, if it's already locked. */
+ FlockDataDirLockFile(fd, filename);
+
pgstat_report_wait_start(WAIT_EVENT_LOCK_FILE_CREATE_READ);
if ((len = read(fd, buffer, sizeof(buffer) - 1)) < 0)
ereport(FATAL,
diff --git a/src/include/pg_config.h.in b/src/include/pg_config.h.in
index 339268dc8ef..e7b8e023829 100644
--- a/src/include/pg_config.h.in
+++ b/src/include/pg_config.h.in
@@ -80,6 +80,10 @@
don't. */
#undef HAVE_DECL_F_FULLFSYNC
+/* Define to 1 if you have the declaration of `F_OFD_SETLK', and to 0 if you
+ don't. */
+#undef HAVE_DECL_F_OFD_SETLK
+
/* Define to 1 if you have the declaration of `memset_s', and to 0 if you
don't. */
#undef HAVE_DECL_MEMSET_S
base-commit: 6831cd9e3b082d7b830c3196742dd49e3540c49b
--
2.49.0
^ permalink raw reply [nested|flat] 10+ messages in thread
* Re: File locks for data directory lockfile in the context of Linux namespaces
2025-12-19 14:27 File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
2026-01-17 15:26 ` Re: File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
@ 2026-06-05 12:37 ` Ilmar Yunusov <tanswis42@gmail.com>
2026-06-19 15:11 ` Re: File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 1 reply; 10+ messages in thread
From: Ilmar Yunusov @ 2026-06-05 12:37 UTC (permalink / raw)
To: pgsql-hackers@lists.postgresql.org; +Cc: Dmitry Dolgov <9erthalion6@gmail.com>
The following review has been posted through the commitfest application:
make installcheck-world: not tested
Implements feature: tested, passed
Spec compliant: not tested
Documentation: not tested
Hi,
I looked at v2, focusing on apply/build status and the PID namespace scenario
described in the cover letter.
I used the v2 patch from Dmitry's 2026-01-17 message, on origin/master at
4cb2a9863d89b320f37eb1bd76822f6f65e59311.
The patch applies cleanly with git am, and git diff --check reports no issues.
I built locally with:
./configure --prefix="$PWD/pg-install" --without-readline --without-zlib --without-icu
make -s -j8
make -s install
The build completed successfully. Configure found F_OFD_SETLK on this macOS
host too.
I also built and tested on Linux with the same configure options:
make -s -j3
make -s install
There, configure also found F_OFD_SETLK, and:
make -C src/test/regress check
passed; all regression tests passed.
For the namespace behavior, I used two PostgreSQL builds on the same Linux host:
unpatched master and v2. The test starts the first postmaster in one
pid/ipc/net namespace, then tries to start a second postmaster in another
pid/ipc/net namespace on the same data directory.
On unpatched master, both postmasters started. Both saw themselves as PID 2 in
their own PID namespace, and the second postmaster rewrote postmaster.pid. After
the second startup, the lock file contained the second socket directory:
2
.../pidns-base/data
1780661958
65437
.../pidns-base/sock2
With v2, the first postmaster started, and the second startup failed with:
FATAL: cannot lock the lock file "postmaster.pid"
HINT: Another server is starting.
So the patch fixes the concrete Linux PID namespace failure mode I tested.
One behavior I wanted to ask about: v2 also OFD-locks Unix socket lock files,
not only the data directory lockfile. That follows from calling
FlockDataDirLockFile() from the generic CreateLockFile() path, and I also saw
it at runtime. With one v2 postmaster using a Unix socket directory, /proc/locks
showed OFD write locks for both files:
.../data/postmaster.pid
.../sock/.s.PGSQL.65436.lock
The corresponding postmaster fds were:
/proc/835058/fd/5 -> .../data/postmaster.pid
/proc/835058/fd/8 -> .../sock/.s.PGSQL.65436.lock
I had expected the new lock to apply only to postmaster.pid, because the patch
title, commit message, helper name, and comments all describe the data directory
lockfile. The socket lockfile behavior therefore looked like either an
intentional scope expansion that should be named as such, or an accidental side
effect of using the generic CreateLockFile() path.
Is locking the socket lock file intentional here? If so, maybe the helper name,
comment, and fd tracking should reflect that broader scope. If not, perhaps the
new lock should be applied only for the isDDLock case; otherwise v2 changes
socket lockfile semantics too, not only postmaster.pid.
While reading that code, I also noticed a small error-path issue: DataDirLockFD
starts as 0, but if fcntl(F_OFD_SETLK) fails with a non-EAGAIN error the code
only emits a warning and does not assign a duplicate fd. UnlinkLockFiles()
then closes DataDirLockFD unconditionally. Initializing it to -1, closing it
conditionally, and checking dup(fd) would make that path more explicit.
Aside from those questions, this looks like a useful best-effort improvement to
me, and the namespace failure mode appears to be addressed by the OFD lock in
the Linux test above.
I did not test NFS behavior, older stable branches, Windows behavior, or
installcheck-world. Unprivileged namespace creation was not permitted on the
Linux host I used, so the namespace repro was run with sudo.
Regards,
Ilmar Yunusov
The new status of this patch is: Waiting on Author
^ permalink raw reply [nested|flat] 10+ messages in thread
* Re: File locks for data directory lockfile in the context of Linux namespaces
2025-12-19 14:27 File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
2026-01-17 15:26 ` Re: File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
2026-06-05 12:37 ` Re: File locks for data directory lockfile in the context of Linux namespaces Ilmar Yunusov <tanswis42@gmail.com>
@ 2026-06-19 15:11 ` Dmitry Dolgov <9erthalion6@gmail.com>
2026-06-23 14:29 ` Re: File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 1 reply; 10+ messages in thread
From: Dmitry Dolgov @ 2026-06-19 15:11 UTC (permalink / raw)
To: Ilmar Yunusov <tanswis42@gmail.com>; +Cc: pgsql-hackers@lists.postgresql.org
> On Fri, Jun 05, 2026 at 12:37:26PM +0000, Ilmar Yunusov wrote:
> The following review has been posted through the commitfest application:
> make installcheck-world: not tested
> Implements feature: tested, passed
> Spec compliant: not tested
> Documentation: not tested
>
> Hi,
>
> I looked at v2, focusing on apply/build status and the PID namespace scenario
> described in the cover letter.
Thanks for the review and testing!
> Is locking the socket lock file intentional here? If so, maybe the helper name,
> comment, and fd tracking should reflect that broader scope. If not, perhaps the
> new lock should be applied only for the isDDLock case; otherwise v2 changes
> socket lockfile semantics too, not only postmaster.pid.
It's been a while, quite frankly I don't remember. On the face of it, I
think both the directory and socket lock files should be locked in the
same way, since both are equally susceptible for the failure scenario
described in this thread, even if the consequences of a failure are
different. Let me do some renaming to clarify that.
> While reading that code, I also noticed a small error-path issue: DataDirLockFD
> starts as 0, but if fcntl(F_OFD_SETLK) fails with a non-EAGAIN error the code
> only emits a warning and does not assign a duplicate fd. UnlinkLockFiles()
> then closes DataDirLockFD unconditionally. Initializing it to -1, closing it
> conditionally, and checking dup(fd) would make that path more explicit.
Good point, I'll address this in the next version.
^ permalink raw reply [nested|flat] 10+ messages in thread
* Re: File locks for data directory lockfile in the context of Linux namespaces
2025-12-19 14:27 File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
2026-01-17 15:26 ` Re: File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
2026-06-05 12:37 ` Re: File locks for data directory lockfile in the context of Linux namespaces Ilmar Yunusov <tanswis42@gmail.com>
2026-06-19 15:11 ` Re: File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
@ 2026-06-23 14:29 ` Dmitry Dolgov <9erthalion6@gmail.com>
2026-07-06 07:30 ` Re: File locks for data directory lockfile in the context of Linux namespaces Ilmar Yunusov <tanswis42@gmail.com>
0 siblings, 1 reply; 10+ messages in thread
From: Dmitry Dolgov @ 2026-06-23 14:29 UTC (permalink / raw)
To: Ilmar Yunusov <tanswis42@gmail.com>; +Cc: pgsql-hackers@lists.postgresql.org
> On Fri, Jun 19, 2026 at 05:11:11PM +0200, Dmitry Dolgov wrote:
> > Is locking the socket lock file intentional here? If so, maybe the helper name,
> > comment, and fd tracking should reflect that broader scope. If not, perhaps the
> > new lock should be applied only for the isDDLock case; otherwise v2 changes
> > socket lockfile semantics too, not only postmaster.pid.
>
> It's been a while, quite frankly I don't remember. On the face of it, I
> think both the directory and socket lock files should be locked in the
> same way, since both are equally susceptible for the failure scenario
> described in this thread, even if the consequences of a failure are
> different. Let me do some renaming to clarify that.
>
> > While reading that code, I also noticed a small error-path issue: DataDirLockFD
> > starts as 0, but if fcntl(F_OFD_SETLK) fails with a non-EAGAIN error the code
> > only emits a warning and does not assign a duplicate fd. UnlinkLockFiles()
> > then closes DataDirLockFD unconditionally. Initializing it to -1, closing it
> > conditionally, and checking dup(fd) would make that path more explicit.
>
> Good point, I'll address this in the next version.
Something like this should be sufficient I think.
From 6301d334002036aa3ddf97236b69ad58eace9e56 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Thu, 18 Dec 2025 18:21:59 +0100
Subject: [PATCH v3] Use open file description locks for lockfiles
When starting up, postmaster checks for an existing data directory lockfile. If
this file contains current process PID, it's assumed to be stale. Turns out
there is another possibility: we might be running in a PID namespace, and there
is another postgres running inside another PID namespace using the same data
directory. The result is that we don't see another process due to namespace
isolation and start concurrently with the other.
To prevent such situations, at startup use fcntl to get an exclusive open file
description lock for data directory lockfile. Since such locks are associated
with open file descriptors, meaning they're not affected by PID namespace
isolation. It's a "best effort" locking, intended to work with already existing
mechanism, not replace it.
This approach was discussed multiple times in the past, and usually was
rejected as the main work horse for the data directory lockfile due to:
* Portability issues. Open file description lock was a non-POSIX extension in
Linux and similar flock is from BSD standard. But looks like everybody agrees
that such locks make more sense than a typical advisory locks, and
F_OFD_SETLK made its way into POSIX.1 2024 [1].
* Issues with NFS. The current state of things here looks like this:
- NFSv3 doesn't implement open file description locks, they're converted to
advisory locks instead. Advisory locks are subject to namespace isolation,
meaning that processes in different PID namespaces will not see each other
advisory lock, and it's still possible to run multiple postgres
instances on the same data directory.
- NFSv4 uses a lease system for locking, I haven't found any mention of
conversion to advisory locks neither in the man page nor in RFC [2].
To summarize, the approach is now considered POSIX and should fix the described
problem everywhere, except NFSv3.
Use open file description lock for both data directory and socker
lockfiles, since both are affected in the same way.
[1]: https://pubs.opengroup.org/onlinepubs/9799919799/functions/fcntl.html
[2]: https://www.rfc-editor.org/rfc/rfc7530
Reviewed-by: Ilmar Yunusov <tanswis42@gmail.com>
---
configure | 14 +++
configure.ac | 3 +
meson.build | 1 +
src/backend/utils/init/miscinit.c | 136 ++++++++++++++++++++++++------
src/include/pg_config.h.in | 4 +
src/tools/pgindent/typedefs.list | 1 +
6 files changed, 134 insertions(+), 25 deletions(-)
diff --git a/configure b/configure
index 5f77f3cac29..15bda6c6413 100755
--- a/configure
+++ b/configure
@@ -16444,6 +16444,20 @@ cat >>confdefs.h <<_ACEOF
_ACEOF
+# Linux open file descriptor locks
+ac_fn_c_check_decl "$LINENO" "F_OFD_SETLK" "ac_cv_have_decl_F_OFD_SETLK" "#include <fcntl.h>
+"
+if test "x$ac_cv_have_decl_F_OFD_SETLK" = xyes; then :
+ ac_have_decl=1
+else
+ ac_have_decl=0
+fi
+
+cat >>confdefs.h <<_ACEOF
+#define HAVE_DECL_F_OFD_SETLK $ac_have_decl
+_ACEOF
+
+
ac_fn_c_check_func "$LINENO" "explicit_bzero" "ac_cv_func_explicit_bzero"
if test "x$ac_cv_func_explicit_bzero" = xyes; then :
$as_echo "#define HAVE_EXPLICIT_BZERO 1" >>confdefs.h
diff --git a/configure.ac b/configure.ac
index 61cee42daa7..e24cb7a6f01 100644
--- a/configure.ac
+++ b/configure.ac
@@ -1913,6 +1913,9 @@ AC_CHECK_DECLS([memset_s], [], [], [#define __STDC_WANT_LIB_EXT1__ 1
# This is probably only present on macOS, but may as well check always
AC_CHECK_DECLS(F_FULLFSYNC, [], [], [#include <fcntl.h>])
+# Linux open file descriptor locks
+AC_CHECK_DECLS([F_OFD_SETLK], [], [], [#include <fcntl.h>])
+
AC_REPLACE_FUNCS(m4_normalize([
explicit_bzero
getopt
diff --git a/meson.build b/meson.build
index 568e0e150bf..3d641dc0403 100644
--- a/meson.build
+++ b/meson.build
@@ -2901,6 +2901,7 @@ decl_checks = [
['strlcpy', 'string.h'],
['strsep', 'string.h'],
['timingsafe_bcmp', 'string.h'],
+ ['F_OFD_SETLK', 'fcntl.h'],
]
# Need to check for function declarations for these functions, because
diff --git a/src/backend/utils/init/miscinit.c b/src/backend/utils/init/miscinit.c
index 7ffc808073a..26c3324542c 100644
--- a/src/backend/utils/init/miscinit.c
+++ b/src/backend/utils/init/miscinit.c
@@ -69,6 +69,15 @@ static List *lock_files = NIL;
static Latch LocalLatchData;
+typedef struct
+{
+ /* LockFile name. */
+ const char *filename;
+
+ /* File descriptor for open file description lock. */
+ int fd;
+} LockFile;
+
/* ----------------------------------------------------------------
* ignoring system indexes support stuff
*
@@ -1119,6 +1128,48 @@ RestoreClientConnectionInfo(char *conninfo)
*-------------------------------------------------------------------------
*/
+/*
+ * OFD lock the specified lockfile.
+ *
+ * Lock the lockfile with an open file description lock. If the lock is already
+ * taken, it's a hard stop. It's only a best effort test, and any other errors
+ * are ignored. On succes the file descriptor is duplicated, to make sure there
+ * will be at least one open copy of it to keep the lock.
+ *
+ * filename argument is used only for reporting purposes.
+ */
+static int
+OFDLockFile(int fd, const char *filename)
+{
+#if HAVE_DECL_F_OFD_SETLK
+ struct flock lock;
+
+ lock.l_type = F_WRLCK;
+ lock.l_whence = SEEK_SET;
+ lock.l_start = 0;
+ lock.l_len = 0;
+ lock.l_pid = 0;
+
+ if (fcntl(fd, F_OFD_SETLK, &lock) == -1)
+ {
+ if (errno == EAGAIN)
+ ereport(FATAL,
+ (errcode(ERRCODE_LOCK_FILE_EXISTS),
+ errmsg("cannot lock the lock file \"%s\"", filename),
+ errhint("Another server is starting.")));
+ else
+ {
+ elog(WARNING, "Failed locking file \"%s\", %m", filename);
+ return -1;
+ }
+ }
+ else
+ return dup(fd);
+#else
+ return -1
+#endif
+}
+
/*
* proc_exit callback to remove lockfiles.
*/
@@ -1129,9 +1180,16 @@ UnlinkLockFiles(int status, Datum arg)
foreach(l, lock_files)
{
- char *curfile = (char *) lfirst(l);
+ LockFile *lock_file = (LockFile *) lfirst(l);
- unlink(curfile);
+ /*
+ * Close the file descriptor, which keeps the open file description
+ * lock.
+ */
+ if (lock_file->fd > 0)
+ close(lock_file->fd);
+
+ unlink(lock_file->filename);
/* Should we complain if the unlink fails? */
}
/* Since we're about to exit, no need to reclaim storage */
@@ -1161,7 +1219,9 @@ CreateLockFile(const char *filename, bool amPostmaster,
const char *socketDir,
bool isDDLock, const char *refName)
{
- int fd;
+ int fd,
+ flock_fd = -1;
+ LockFile *lock_file;
char buffer[MAXPGPATH * 2 + 256];
int ntries;
int len;
@@ -1173,22 +1233,32 @@ CreateLockFile(const char *filename, bool amPostmaster,
const char *envvar;
/*
- * If the PID in the lockfile is our own PID or our parent's or
- * grandparent's PID, then the file must be stale (probably left over from
- * a previous system boot cycle). We need to check this because of the
- * likelihood that a reboot will assign exactly the same PID as we had in
- * the previous reboot, or one that's only one or two counts larger and
- * hence the lockfile's PID now refers to an ancestor shell process. We
- * allow pg_ctl to pass down its parent shell PID (our grandparent PID)
- * via the environment variable PG_GRANDPARENT_PID; this is so that
- * launching the postmaster via pg_ctl can be just as reliable as
- * launching it directly. There is no provision for detecting
- * further-removed ancestor processes, but if the init script is written
- * carefully then all but the immediate parent shell will be root-owned
- * processes and so the kill test will fail with EPERM. Note that we
- * cannot get a false negative this way, because an existing postmaster
- * would surely never launch a competing postmaster or pg_ctl process
- * directly.
+ * If we find an already existing lockfile containing our own PID, there
+ * are few options:
+ *
+ * - There is another process, that we don't see due to PID namespace
+ * isolation, which is already running in this data directory.
+ *
+ * To prevent two concurrent processes working with the same data
+ * directory, we first try to lock the lockfile exclusively.
+ *
+ * - The file must be stale, probably left over from a previous system
+ * boot cycle. The same if the lockfile contains our parent's or
+ * grandparent's PID.
+ *
+ * We need to check this because of the likelihood that a reboot will
+ * assign exactly the same PID as we had in the previous reboot, or one
+ * that's only one or two counts larger and hence the lockfile's PID now
+ * refers to an ancestor shell process. We allow pg_ctl to pass down its
+ * parent shell PID (our grandparent PID) via the environment variable
+ * PG_GRANDPARENT_PID; this is so that launching the postmaster via pg_ctl
+ * can be just as reliable as launching it directly. There is no
+ * provision for detecting further-removed ancestor processes, but if the
+ * init script is written carefully then all but the immediate parent
+ * shell will be root-owned processes and so the kill test will fail with
+ * EPERM. Note that we cannot get a false negative this way, because an
+ * existing postmaster would surely never launch a competing postmaster or
+ * pg_ctl process directly.
*/
my_pid = getpid();
@@ -1224,7 +1294,11 @@ CreateLockFile(const char *filename, bool amPostmaster,
*/
fd = open(filename, O_RDWR | O_CREAT | O_EXCL, pg_file_create_mode);
if (fd >= 0)
- break; /* Success; exit the retry loop */
+ {
+ /* Success; lock and exit the retry loop */
+ flock_fd = OFDLockFile(fd, filename);
+ break;
+ }
/*
* Couldn't create the pid file. Probably it already exists.
@@ -1238,8 +1312,12 @@ CreateLockFile(const char *filename, bool amPostmaster,
/*
* Read the file to get the old owner's PID. Note race condition
* here: file might have been deleted since we tried to create it.
+ *
+ * We're going to use the same fd for flock, and want to create a
+ * write lock for the latter one. Since both fd and the lock have to
+ * be of the same type, open the file for read and write.
*/
- fd = open(filename, O_RDONLY, pg_file_create_mode);
+ fd = open(filename, O_RDWR, pg_file_create_mode);
if (fd < 0)
{
if (errno == ENOENT)
@@ -1249,6 +1327,10 @@ CreateLockFile(const char *filename, bool amPostmaster,
errmsg("could not open lock file \"%s\": %m",
filename)));
}
+
+ /* Try to lock the file. We stop here, if it's already locked. */
+ flock_fd = OFDLockFile(fd, filename);
+
pgstat_report_wait_start(WAIT_EVENT_LOCK_FILE_CREATE_READ);
if ((len = read(fd, buffer, sizeof(buffer) - 1)) < 0)
ereport(FATAL,
@@ -1448,7 +1530,11 @@ CreateLockFile(const char *filename, bool amPostmaster,
* Use lcons so that the lock files are unlinked in reverse order of
* creation; this is critical!
*/
- lock_files = lcons(pstrdup(filename), lock_files);
+ lock_file = palloc0_object(LockFile);
+ lock_file->filename = pstrdup(filename);
+ lock_file->fd = flock_fd;
+
+ lock_files = lcons(lock_file, lock_files);
}
/*
@@ -1495,14 +1581,14 @@ TouchSocketLockFiles(void)
foreach(l, lock_files)
{
- char *socketLockFile = (char *) lfirst(l);
+ LockFile *lock_file = (LockFile *) lfirst(l);
/* No need to touch the data directory lock file, we trust */
- if (strcmp(socketLockFile, DIRECTORY_LOCK_FILE) == 0)
+ if (strcmp(lock_file->filename, DIRECTORY_LOCK_FILE) == 0)
continue;
/* we just ignore any error here */
- (void) utime(socketLockFile, NULL);
+ (void) utime(lock_file->filename, NULL);
}
}
diff --git a/src/include/pg_config.h.in b/src/include/pg_config.h.in
index 4f8113c144b..cc38c06dc13 100644
--- a/src/include/pg_config.h.in
+++ b/src/include/pg_config.h.in
@@ -85,6 +85,10 @@
don't. */
#undef HAVE_DECL_F_FULLFSYNC
+/* Define to 1 if you have the declaration of `F_OFD_SETLK', and to 0 if you
+ don't. */
+#undef HAVE_DECL_F_OFD_SETLK
+
/* Define to 1 if you have the declaration of `memset_s', and to 0 if you
don't. */
#undef HAVE_DECL_MEMSET_S
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 1969d467c1d..185d69b5520 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -1653,6 +1653,7 @@ LocationLen
LockAcquireResult
LockClauseStrength
LockData
+LockFile
LockInfoData
LockInstanceData
LockMethod
base-commit: 031904048aa22e7c70dc8e9c170e2743f9b0f090
--
2.52.0
Attachments:
[text/plain] v3-0001-Use-open-file-description-locks-for-lockfiles.patch (12.5K, ../../xhzl7ll7fwmlecg4htsgvr3fcwcec4cvugqebphofvcz4srkkn@vpwa2zgswuwy/2-v3-0001-Use-open-file-description-locks-for-lockfiles.patch)
download | inline diff:
From 6301d334002036aa3ddf97236b69ad58eace9e56 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Thu, 18 Dec 2025 18:21:59 +0100
Subject: [PATCH v3] Use open file description locks for lockfiles
When starting up, postmaster checks for an existing data directory lockfile. If
this file contains current process PID, it's assumed to be stale. Turns out
there is another possibility: we might be running in a PID namespace, and there
is another postgres running inside another PID namespace using the same data
directory. The result is that we don't see another process due to namespace
isolation and start concurrently with the other.
To prevent such situations, at startup use fcntl to get an exclusive open file
description lock for data directory lockfile. Since such locks are associated
with open file descriptors, meaning they're not affected by PID namespace
isolation. It's a "best effort" locking, intended to work with already existing
mechanism, not replace it.
This approach was discussed multiple times in the past, and usually was
rejected as the main work horse for the data directory lockfile due to:
* Portability issues. Open file description lock was a non-POSIX extension in
Linux and similar flock is from BSD standard. But looks like everybody agrees
that such locks make more sense than a typical advisory locks, and
F_OFD_SETLK made its way into POSIX.1 2024 [1].
* Issues with NFS. The current state of things here looks like this:
- NFSv3 doesn't implement open file description locks, they're converted to
advisory locks instead. Advisory locks are subject to namespace isolation,
meaning that processes in different PID namespaces will not see each other
advisory lock, and it's still possible to run multiple postgres
instances on the same data directory.
- NFSv4 uses a lease system for locking, I haven't found any mention of
conversion to advisory locks neither in the man page nor in RFC [2].
To summarize, the approach is now considered POSIX and should fix the described
problem everywhere, except NFSv3.
Use open file description lock for both data directory and socker
lockfiles, since both are affected in the same way.
[1]: https://pubs.opengroup.org/onlinepubs/9799919799/functions/fcntl.html
[2]: https://www.rfc-editor.org/rfc/rfc7530
Reviewed-by: Ilmar Yunusov <tanswis42@gmail.com>
---
configure | 14 +++
configure.ac | 3 +
meson.build | 1 +
src/backend/utils/init/miscinit.c | 136 ++++++++++++++++++++++++------
src/include/pg_config.h.in | 4 +
src/tools/pgindent/typedefs.list | 1 +
6 files changed, 134 insertions(+), 25 deletions(-)
diff --git a/configure b/configure
index 5f77f3cac29..15bda6c6413 100755
--- a/configure
+++ b/configure
@@ -16444,6 +16444,20 @@ cat >>confdefs.h <<_ACEOF
_ACEOF
+# Linux open file descriptor locks
+ac_fn_c_check_decl "$LINENO" "F_OFD_SETLK" "ac_cv_have_decl_F_OFD_SETLK" "#include <fcntl.h>
+"
+if test "x$ac_cv_have_decl_F_OFD_SETLK" = xyes; then :
+ ac_have_decl=1
+else
+ ac_have_decl=0
+fi
+
+cat >>confdefs.h <<_ACEOF
+#define HAVE_DECL_F_OFD_SETLK $ac_have_decl
+_ACEOF
+
+
ac_fn_c_check_func "$LINENO" "explicit_bzero" "ac_cv_func_explicit_bzero"
if test "x$ac_cv_func_explicit_bzero" = xyes; then :
$as_echo "#define HAVE_EXPLICIT_BZERO 1" >>confdefs.h
diff --git a/configure.ac b/configure.ac
index 61cee42daa7..e24cb7a6f01 100644
--- a/configure.ac
+++ b/configure.ac
@@ -1913,6 +1913,9 @@ AC_CHECK_DECLS([memset_s], [], [], [#define __STDC_WANT_LIB_EXT1__ 1
# This is probably only present on macOS, but may as well check always
AC_CHECK_DECLS(F_FULLFSYNC, [], [], [#include <fcntl.h>])
+# Linux open file descriptor locks
+AC_CHECK_DECLS([F_OFD_SETLK], [], [], [#include <fcntl.h>])
+
AC_REPLACE_FUNCS(m4_normalize([
explicit_bzero
getopt
diff --git a/meson.build b/meson.build
index 568e0e150bf..3d641dc0403 100644
--- a/meson.build
+++ b/meson.build
@@ -2901,6 +2901,7 @@ decl_checks = [
['strlcpy', 'string.h'],
['strsep', 'string.h'],
['timingsafe_bcmp', 'string.h'],
+ ['F_OFD_SETLK', 'fcntl.h'],
]
# Need to check for function declarations for these functions, because
diff --git a/src/backend/utils/init/miscinit.c b/src/backend/utils/init/miscinit.c
index 7ffc808073a..26c3324542c 100644
--- a/src/backend/utils/init/miscinit.c
+++ b/src/backend/utils/init/miscinit.c
@@ -69,6 +69,15 @@ static List *lock_files = NIL;
static Latch LocalLatchData;
+typedef struct
+{
+ /* LockFile name. */
+ const char *filename;
+
+ /* File descriptor for open file description lock. */
+ int fd;
+} LockFile;
+
/* ----------------------------------------------------------------
* ignoring system indexes support stuff
*
@@ -1119,6 +1128,48 @@ RestoreClientConnectionInfo(char *conninfo)
*-------------------------------------------------------------------------
*/
+/*
+ * OFD lock the specified lockfile.
+ *
+ * Lock the lockfile with an open file description lock. If the lock is already
+ * taken, it's a hard stop. It's only a best effort test, and any other errors
+ * are ignored. On succes the file descriptor is duplicated, to make sure there
+ * will be at least one open copy of it to keep the lock.
+ *
+ * filename argument is used only for reporting purposes.
+ */
+static int
+OFDLockFile(int fd, const char *filename)
+{
+#if HAVE_DECL_F_OFD_SETLK
+ struct flock lock;
+
+ lock.l_type = F_WRLCK;
+ lock.l_whence = SEEK_SET;
+ lock.l_start = 0;
+ lock.l_len = 0;
+ lock.l_pid = 0;
+
+ if (fcntl(fd, F_OFD_SETLK, &lock) == -1)
+ {
+ if (errno == EAGAIN)
+ ereport(FATAL,
+ (errcode(ERRCODE_LOCK_FILE_EXISTS),
+ errmsg("cannot lock the lock file \"%s\"", filename),
+ errhint("Another server is starting.")));
+ else
+ {
+ elog(WARNING, "Failed locking file \"%s\", %m", filename);
+ return -1;
+ }
+ }
+ else
+ return dup(fd);
+#else
+ return -1
+#endif
+}
+
/*
* proc_exit callback to remove lockfiles.
*/
@@ -1129,9 +1180,16 @@ UnlinkLockFiles(int status, Datum arg)
foreach(l, lock_files)
{
- char *curfile = (char *) lfirst(l);
+ LockFile *lock_file = (LockFile *) lfirst(l);
- unlink(curfile);
+ /*
+ * Close the file descriptor, which keeps the open file description
+ * lock.
+ */
+ if (lock_file->fd > 0)
+ close(lock_file->fd);
+
+ unlink(lock_file->filename);
/* Should we complain if the unlink fails? */
}
/* Since we're about to exit, no need to reclaim storage */
@@ -1161,7 +1219,9 @@ CreateLockFile(const char *filename, bool amPostmaster,
const char *socketDir,
bool isDDLock, const char *refName)
{
- int fd;
+ int fd,
+ flock_fd = -1;
+ LockFile *lock_file;
char buffer[MAXPGPATH * 2 + 256];
int ntries;
int len;
@@ -1173,22 +1233,32 @@ CreateLockFile(const char *filename, bool amPostmaster,
const char *envvar;
/*
- * If the PID in the lockfile is our own PID or our parent's or
- * grandparent's PID, then the file must be stale (probably left over from
- * a previous system boot cycle). We need to check this because of the
- * likelihood that a reboot will assign exactly the same PID as we had in
- * the previous reboot, or one that's only one or two counts larger and
- * hence the lockfile's PID now refers to an ancestor shell process. We
- * allow pg_ctl to pass down its parent shell PID (our grandparent PID)
- * via the environment variable PG_GRANDPARENT_PID; this is so that
- * launching the postmaster via pg_ctl can be just as reliable as
- * launching it directly. There is no provision for detecting
- * further-removed ancestor processes, but if the init script is written
- * carefully then all but the immediate parent shell will be root-owned
- * processes and so the kill test will fail with EPERM. Note that we
- * cannot get a false negative this way, because an existing postmaster
- * would surely never launch a competing postmaster or pg_ctl process
- * directly.
+ * If we find an already existing lockfile containing our own PID, there
+ * are few options:
+ *
+ * - There is another process, that we don't see due to PID namespace
+ * isolation, which is already running in this data directory.
+ *
+ * To prevent two concurrent processes working with the same data
+ * directory, we first try to lock the lockfile exclusively.
+ *
+ * - The file must be stale, probably left over from a previous system
+ * boot cycle. The same if the lockfile contains our parent's or
+ * grandparent's PID.
+ *
+ * We need to check this because of the likelihood that a reboot will
+ * assign exactly the same PID as we had in the previous reboot, or one
+ * that's only one or two counts larger and hence the lockfile's PID now
+ * refers to an ancestor shell process. We allow pg_ctl to pass down its
+ * parent shell PID (our grandparent PID) via the environment variable
+ * PG_GRANDPARENT_PID; this is so that launching the postmaster via pg_ctl
+ * can be just as reliable as launching it directly. There is no
+ * provision for detecting further-removed ancestor processes, but if the
+ * init script is written carefully then all but the immediate parent
+ * shell will be root-owned processes and so the kill test will fail with
+ * EPERM. Note that we cannot get a false negative this way, because an
+ * existing postmaster would surely never launch a competing postmaster or
+ * pg_ctl process directly.
*/
my_pid = getpid();
@@ -1224,7 +1294,11 @@ CreateLockFile(const char *filename, bool amPostmaster,
*/
fd = open(filename, O_RDWR | O_CREAT | O_EXCL, pg_file_create_mode);
if (fd >= 0)
- break; /* Success; exit the retry loop */
+ {
+ /* Success; lock and exit the retry loop */
+ flock_fd = OFDLockFile(fd, filename);
+ break;
+ }
/*
* Couldn't create the pid file. Probably it already exists.
@@ -1238,8 +1312,12 @@ CreateLockFile(const char *filename, bool amPostmaster,
/*
* Read the file to get the old owner's PID. Note race condition
* here: file might have been deleted since we tried to create it.
+ *
+ * We're going to use the same fd for flock, and want to create a
+ * write lock for the latter one. Since both fd and the lock have to
+ * be of the same type, open the file for read and write.
*/
- fd = open(filename, O_RDONLY, pg_file_create_mode);
+ fd = open(filename, O_RDWR, pg_file_create_mode);
if (fd < 0)
{
if (errno == ENOENT)
@@ -1249,6 +1327,10 @@ CreateLockFile(const char *filename, bool amPostmaster,
errmsg("could not open lock file \"%s\": %m",
filename)));
}
+
+ /* Try to lock the file. We stop here, if it's already locked. */
+ flock_fd = OFDLockFile(fd, filename);
+
pgstat_report_wait_start(WAIT_EVENT_LOCK_FILE_CREATE_READ);
if ((len = read(fd, buffer, sizeof(buffer) - 1)) < 0)
ereport(FATAL,
@@ -1448,7 +1530,11 @@ CreateLockFile(const char *filename, bool amPostmaster,
* Use lcons so that the lock files are unlinked in reverse order of
* creation; this is critical!
*/
- lock_files = lcons(pstrdup(filename), lock_files);
+ lock_file = palloc0_object(LockFile);
+ lock_file->filename = pstrdup(filename);
+ lock_file->fd = flock_fd;
+
+ lock_files = lcons(lock_file, lock_files);
}
/*
@@ -1495,14 +1581,14 @@ TouchSocketLockFiles(void)
foreach(l, lock_files)
{
- char *socketLockFile = (char *) lfirst(l);
+ LockFile *lock_file = (LockFile *) lfirst(l);
/* No need to touch the data directory lock file, we trust */
- if (strcmp(socketLockFile, DIRECTORY_LOCK_FILE) == 0)
+ if (strcmp(lock_file->filename, DIRECTORY_LOCK_FILE) == 0)
continue;
/* we just ignore any error here */
- (void) utime(socketLockFile, NULL);
+ (void) utime(lock_file->filename, NULL);
}
}
diff --git a/src/include/pg_config.h.in b/src/include/pg_config.h.in
index 4f8113c144b..cc38c06dc13 100644
--- a/src/include/pg_config.h.in
+++ b/src/include/pg_config.h.in
@@ -85,6 +85,10 @@
don't. */
#undef HAVE_DECL_F_FULLFSYNC
+/* Define to 1 if you have the declaration of `F_OFD_SETLK', and to 0 if you
+ don't. */
+#undef HAVE_DECL_F_OFD_SETLK
+
/* Define to 1 if you have the declaration of `memset_s', and to 0 if you
don't. */
#undef HAVE_DECL_MEMSET_S
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 1969d467c1d..185d69b5520 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -1653,6 +1653,7 @@ LocationLen
LockAcquireResult
LockClauseStrength
LockData
+LockFile
LockInfoData
LockInstanceData
LockMethod
base-commit: 031904048aa22e7c70dc8e9c170e2743f9b0f090
--
2.52.0
^ permalink raw reply [nested|flat] 10+ messages in thread
* Re: File locks for data directory lockfile in the context of Linux namespaces
2025-12-19 14:27 File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
2026-01-17 15:26 ` Re: File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
2026-06-05 12:37 ` Re: File locks for data directory lockfile in the context of Linux namespaces Ilmar Yunusov <tanswis42@gmail.com>
2026-06-19 15:11 ` Re: File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
2026-06-23 14:29 ` Re: File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
@ 2026-07-06 07:30 ` Ilmar Yunusov <tanswis42@gmail.com>
2026-07-06 12:10 ` Re: File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 1 reply; 10+ messages in thread
From: Ilmar Yunusov @ 2026-07-06 07:30 UTC (permalink / raw)
To: pgsql-hackers@lists.postgresql.org; +Cc: Dmitry Dolgov <erthalion.new@posteo.com>
The following review has been posted through the commitfest application:
make installcheck-world: not tested
Implements feature: tested, passed
Spec compliant: not tested
Documentation: not tested
Hi,
I looked at v2, focusing on apply/build status and the PID namespace scenario
described in the cover letter.
I used the v2 patch from Dmitry's 2026-01-17 message, on origin/master at
4cb2a9863d89b320f37eb1bd76822f6f65e59311.
The patch applies cleanly with git am, and git diff --check reports no issues.
I built locally with:
./configure --prefix="$PWD/pg-install" --without-readline --without-zlib --without-icu
make -s -j8
make -s install
The build completed successfully. Configure found F_OFD_SETLK on this macOS
host too.
I also built and tested on Linux with the same configure options:
make -s -j3
make -s install
There, configure also found F_OFD_SETLK, and:
make -C src/test/regress check
passed; all regression tests passed.
For the namespace behavior, I used two PostgreSQL builds on the same Linux host:
unpatched master and v2. The test starts the first postmaster in one
pid/ipc/net namespace, then tries to start a second postmaster in another
pid/ipc/net namespace on the same data directory.
On unpatched master, both postmasters started. Both saw themselves as PID 2 in
their own PID namespace, and the second postmaster rewrote postmaster.pid. After
the second startup, the lock file contained the second socket directory:
2
.../pidns-base/data
1780661958
65437
.../pidns-base/sock2
With v2, the first postmaster started, and the second startup failed with:
FATAL: cannot lock the lock file "postmaster.pid"
HINT: Another server is starting.
So the patch fixes the concrete Linux PID namespace failure mode I tested.
One behavior I wanted to ask about: v2 also OFD-locks Unix socket lock files,
not only the data directory lockfile. That follows from calling
FlockDataDirLockFile() from the generic CreateLockFile() path, and I also saw
it at runtime. With one v2 postmaster using a Unix socket directory, /proc/locks
showed OFD write locks for both files:
.../data/postmaster.pid
.../sock/.s.PGSQL.65436.lock
The corresponding postmaster fds were:
/proc/835058/fd/5 -> .../data/postmaster.pid
/proc/835058/fd/8 -> .../sock/.s.PGSQL.65436.lock
I had expected the new lock to apply only to postmaster.pid, because the patch
title, commit message, helper name, and comments all describe the data directory
lockfile. The socket lockfile behavior therefore looked like either an
intentional scope expansion that should be named as such, or an accidental side
effect of using the generic CreateLockFile() path.
Is locking the socket lock file intentional here? If so, maybe the helper name,
comment, and fd tracking should reflect that broader scope. If not, perhaps the
new lock should be applied only for the isDDLock case; otherwise v2 changes
socket lockfile semantics too, not only postmaster.pid.
While reading that code, I also noticed a small error-path issue: DataDirLockFD
starts as 0, but if fcntl(F_OFD_SETLK) fails with a non-EAGAIN error the code
only emits a warning and does not assign a duplicate fd. UnlinkLockFiles()
then closes DataDirLockFD unconditionally. Initializing it to -1, closing it
conditionally, and checking dup(fd) would make that path more explicit.
Aside from those questions, this looks like a useful best-effort improvement to
me, and the namespace failure mode appears to be addressed by the OFD lock in
the Linux test above.
I did not test NFS behavior, older stable branches, Windows behavior, or
installcheck-world. Unprivileged namespace creation was not permitted on the
Linux host I used, so the namespace repro was run with sudo.
Regards,
Ilmar Yunusov
## v3 follow-up draft
Suggested CommitFest form:
- In response to: Dmitry's 2026-06-23 v3 message
- make installcheck-world: not tested
- Implements feature: tested, passed
- Spec compliant: not tested
- Documentation: not tested
- New status: Waiting on Author
- Attachments: none
Subject:
Re: File locks for data directory lockfile in the context of Linux namespaces
Body:
Hi,
I looked at v3, focusing on whether it addresses the two points from my
previous review, the Linux runtime behavior, and the current CFBot failures.
I used the v3 patch from Dmitry's 2026-06-23 message, on origin/master at
9d1188f29865e66c4196578501e74e8c815fba8d.
The patch applies cleanly with git am, and git diff --check reports no issues.
On Ubuntu 24.04.4, configure found F_OFD_SETLK, and this passed:
./configure --prefix=/home/master/pgcf-6335-v3-20260706/pg-install --without-readline --without-zlib --without-icu
make -s -j3
make -s install
make -C src/test/regress check
All 245 regression tests passed.
I re-ran the PID namespace reproducer from the earlier review against v3. The
first postmaster started in one pid/ipc/net namespace on the test data
directory. A second postmaster in another pid/ipc/net namespace on the same
data directory failed with:
FATAL: cannot lock the lock file "postmaster.pid"
HINT: Another server is starting.
So the concrete Linux PID namespace failure mode I tested still looks addressed
in v3.
I also checked the lockfile scope at runtime. With one v3 postmaster using a
Unix socket directory, /proc/<pid>/fd showed fds for both files:
.../data/postmaster.pid
.../sock/.s.PGSQL.65442.lock
/proc/locks also showed OFDLCK ADVISORY WRITE entries for both file inodes.
That matches the v3 scope where socket lockfiles are intentionally locked too.
The other question from my previous review also looks addressed: the single
static DataDirLockFD issue is replaced by per-lockfile fd tracking in the
lock_files list, with conditional close in UnlinkLockFiles().
There is still a build blocker, matching the current CFBot failures. v3 adds a
typedef named LockFile in src/backend/utils/init/miscinit.c, but that name
conflicts with the Windows LockFile() function. The CompilerWarnings log
reports, for example:
miscinit.c:79:3: error: 'LockFile' redeclared as different kind of symbol
and the Windows logs show the same issue:
miscinit.c(79): error C2365: 'LockFile': redefinition; previous definition was 'function'
There is also a missing semicolon in the non-F_OFD_SETLK branch:
miscinit.c:1169:18: error: expected ';' before '}' token
As a quick local compile check, in the temporary worktree only, renaming the
new typedef to LockFileData and adding the missing semicolon made the normal
macOS build pass. With HAVE_DECL_F_OFD_SETLK temporarily forced to 0 in the
generated pg_config.h, the affected miscinit.o target also compiled after those
two mechanical changes. I did not run a local Windows build.
I did not run installcheck-world, and did not test NFS behavior or
stable-branch backpatch applicability.
Regards,
Ilmar Yunusov
The new status of this patch is: Waiting on Author
^ permalink raw reply [nested|flat] 10+ messages in thread
* Re: File locks for data directory lockfile in the context of Linux namespaces
2025-12-19 14:27 File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
2026-01-17 15:26 ` Re: File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
2026-06-05 12:37 ` Re: File locks for data directory lockfile in the context of Linux namespaces Ilmar Yunusov <tanswis42@gmail.com>
2026-06-19 15:11 ` Re: File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
2026-06-23 14:29 ` Re: File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
2026-07-06 07:30 ` Re: File locks for data directory lockfile in the context of Linux namespaces Ilmar Yunusov <tanswis42@gmail.com>
@ 2026-07-06 12:10 ` Dmitry Dolgov <9erthalion6@gmail.com>
2026-07-06 22:07 ` Re: File locks for data directory lockfile in the context of Linux namespaces Zsolt Parragi <zsolt.parragi@percona.com>
0 siblings, 1 reply; 10+ messages in thread
From: Dmitry Dolgov @ 2026-07-06 12:10 UTC (permalink / raw)
To: Ilmar Yunusov <tanswis42@gmail.com>; +Cc: pgsql-hackers@lists.postgresql.org
> On Mon, Jul 06, 2026 at 07:30:04AM +0000, Ilmar Yunusov wrote:
>
> There is still a build blocker, matching the current CFBot failures. v3 adds a
> typedef named LockFile in src/backend/utils/init/miscinit.c, but that name
> conflicts with the Windows LockFile() function. The CompilerWarnings log
> reports, for example:
Thanks for looking into the build failure. I wanted to check it out what
was happening on Windows, but after the migration from Cirrus to Github
one have to be logged in to see the build logs, and at that moment I
found myself logged off.
Regarding the name, I'm afraid LockFileData has a chance of causing some
confusion due to a common pattern, where structures are named with
"Data" suffix and a pointer type definition without. Since it's about a
name clash with an external library, I suggest PGLockFile instead, but
open for better suggestions.
From 2ebad2531437fb1a943e7337981098c1268d087a Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Thu, 18 Dec 2025 18:21:59 +0100
Subject: [PATCH v4] Use open file description locks for lockfiles
When starting up, postmaster checks for an existing data directory lockfile. If
this file contains current process PID, it's assumed to be stale. Turns out
there is another possibility: we might be running in a PID namespace, and there
is another postgres running inside another PID namespace using the same data
directory. The result is that we don't see another process due to namespace
isolation and start concurrently with the other.
To prevent such situations, at startup use fcntl to get an exclusive open file
description lock for data directory lockfile. Since such locks are associated
with open file descriptors, meaning they're not affected by PID namespace
isolation. It's a "best effort" locking, intended to work with already existing
mechanism, not replace it.
This approach was discussed multiple times in the past, and usually was
rejected as the main work horse for the data directory lockfile due to:
* Portability issues. Open file description lock was a non-POSIX extension in
Linux and similar flock is from BSD standard. But looks like everybody agrees
that such locks make more sense than a typical advisory locks, and
F_OFD_SETLK made its way into POSIX.1 2024 [1].
* Issues with NFS. The current state of things here looks like this:
- NFSv3 doesn't implement open file description locks, they're converted to
advisory locks instead. Advisory locks are subject to namespace isolation,
meaning that processes in different PID namespaces will not see each other
advisory lock, and it's still possible to run multiple postgres
instances on the same data directory.
- NFSv4 uses a lease system for locking, I haven't found any mention of
conversion to advisory locks neither in the man page nor in RFC [2].
To summarize, the approach is now considered POSIX and should fix the described
problem everywhere, except NFSv3.
Use open file description lock for both data directory and socker
lockfiles, since both are affected in the same way.
[1]: https://pubs.opengroup.org/onlinepubs/9799919799/functions/fcntl.html
[2]: https://www.rfc-editor.org/rfc/rfc7530
Reviewed-by: Ilmar Yunusov <tanswis42@gmail.com>
---
configure | 14 +++
configure.ac | 3 +
meson.build | 1 +
src/backend/utils/init/miscinit.c | 136 ++++++++++++++++++++++++------
src/include/pg_config.h.in | 4 +
src/tools/pgindent/typedefs.list | 1 +
6 files changed, 134 insertions(+), 25 deletions(-)
diff --git a/configure b/configure
index 35b0b72f0a7..25ebcd3cc47 100755
--- a/configure
+++ b/configure
@@ -16444,6 +16444,20 @@ cat >>confdefs.h <<_ACEOF
_ACEOF
+# Linux open file descriptor locks
+ac_fn_c_check_decl "$LINENO" "F_OFD_SETLK" "ac_cv_have_decl_F_OFD_SETLK" "#include <fcntl.h>
+"
+if test "x$ac_cv_have_decl_F_OFD_SETLK" = xyes; then :
+ ac_have_decl=1
+else
+ ac_have_decl=0
+fi
+
+cat >>confdefs.h <<_ACEOF
+#define HAVE_DECL_F_OFD_SETLK $ac_have_decl
+_ACEOF
+
+
ac_fn_c_check_func "$LINENO" "explicit_bzero" "ac_cv_func_explicit_bzero"
if test "x$ac_cv_func_explicit_bzero" = xyes; then :
$as_echo "#define HAVE_EXPLICIT_BZERO 1" >>confdefs.h
diff --git a/configure.ac b/configure.ac
index 0e624fe36b9..677137207e7 100644
--- a/configure.ac
+++ b/configure.ac
@@ -1913,6 +1913,9 @@ AC_CHECK_DECLS([memset_s], [], [], [#define __STDC_WANT_LIB_EXT1__ 1
# This is probably only present on macOS, but may as well check always
AC_CHECK_DECLS(F_FULLFSYNC, [], [], [#include <fcntl.h>])
+# Linux open file descriptor locks
+AC_CHECK_DECLS([F_OFD_SETLK], [], [], [#include <fcntl.h>])
+
AC_REPLACE_FUNCS(m4_normalize([
explicit_bzero
getopt
diff --git a/meson.build b/meson.build
index d88a7a70308..153fbb477bb 100644
--- a/meson.build
+++ b/meson.build
@@ -2901,6 +2901,7 @@ decl_checks = [
['strlcpy', 'string.h'],
['strsep', 'string.h'],
['timingsafe_bcmp', 'string.h'],
+ ['F_OFD_SETLK', 'fcntl.h'],
]
# Need to check for function declarations for these functions, because
diff --git a/src/backend/utils/init/miscinit.c b/src/backend/utils/init/miscinit.c
index 7ffc808073a..d4c2f80eb46 100644
--- a/src/backend/utils/init/miscinit.c
+++ b/src/backend/utils/init/miscinit.c
@@ -69,6 +69,15 @@ static List *lock_files = NIL;
static Latch LocalLatchData;
+typedef struct
+{
+ /* LockFile name. */
+ const char *filename;
+
+ /* File descriptor for open file description lock. */
+ int fd;
+} PGLockFile;
+
/* ----------------------------------------------------------------
* ignoring system indexes support stuff
*
@@ -1119,6 +1128,48 @@ RestoreClientConnectionInfo(char *conninfo)
*-------------------------------------------------------------------------
*/
+/*
+ * OFD lock the specified lockfile.
+ *
+ * Lock the lockfile with an open file description lock. If the lock is already
+ * taken, it's a hard stop. It's only a best effort test, and any other errors
+ * are ignored. On succes the file descriptor is duplicated, to make sure there
+ * will be at least one open copy of it to keep the lock.
+ *
+ * filename argument is used only for reporting purposes.
+ */
+static int
+OFDLockFile(int fd, const char *filename)
+{
+#if HAVE_DECL_F_OFD_SETLK
+ struct flock lock;
+
+ lock.l_type = F_WRLCK;
+ lock.l_whence = SEEK_SET;
+ lock.l_start = 0;
+ lock.l_len = 0;
+ lock.l_pid = 0;
+
+ if (fcntl(fd, F_OFD_SETLK, &lock) == -1)
+ {
+ if (errno == EAGAIN)
+ ereport(FATAL,
+ (errcode(ERRCODE_LOCK_FILE_EXISTS),
+ errmsg("cannot lock the lock file \"%s\"", filename),
+ errhint("Another server is starting.")));
+ else
+ {
+ elog(WARNING, "Failed locking file \"%s\", %m", filename);
+ return -1;
+ }
+ }
+ else
+ return dup(fd);
+#else
+ return -1;
+#endif
+}
+
/*
* proc_exit callback to remove lockfiles.
*/
@@ -1129,9 +1180,16 @@ UnlinkLockFiles(int status, Datum arg)
foreach(l, lock_files)
{
- char *curfile = (char *) lfirst(l);
+ PGLockFile *lock_file = (PGLockFile *) lfirst(l);
- unlink(curfile);
+ /*
+ * Close the file descriptor, which keeps the open file description
+ * lock.
+ */
+ if (lock_file->fd > 0)
+ close(lock_file->fd);
+
+ unlink(lock_file->filename);
/* Should we complain if the unlink fails? */
}
/* Since we're about to exit, no need to reclaim storage */
@@ -1161,7 +1219,9 @@ CreateLockFile(const char *filename, bool amPostmaster,
const char *socketDir,
bool isDDLock, const char *refName)
{
- int fd;
+ int fd,
+ flock_fd = -1;
+ PGLockFile *lock_file;
char buffer[MAXPGPATH * 2 + 256];
int ntries;
int len;
@@ -1173,22 +1233,32 @@ CreateLockFile(const char *filename, bool amPostmaster,
const char *envvar;
/*
- * If the PID in the lockfile is our own PID or our parent's or
- * grandparent's PID, then the file must be stale (probably left over from
- * a previous system boot cycle). We need to check this because of the
- * likelihood that a reboot will assign exactly the same PID as we had in
- * the previous reboot, or one that's only one or two counts larger and
- * hence the lockfile's PID now refers to an ancestor shell process. We
- * allow pg_ctl to pass down its parent shell PID (our grandparent PID)
- * via the environment variable PG_GRANDPARENT_PID; this is so that
- * launching the postmaster via pg_ctl can be just as reliable as
- * launching it directly. There is no provision for detecting
- * further-removed ancestor processes, but if the init script is written
- * carefully then all but the immediate parent shell will be root-owned
- * processes and so the kill test will fail with EPERM. Note that we
- * cannot get a false negative this way, because an existing postmaster
- * would surely never launch a competing postmaster or pg_ctl process
- * directly.
+ * If we find an already existing lockfile containing our own PID, there
+ * are few options:
+ *
+ * - There is another process, that we don't see due to PID namespace
+ * isolation, which is already running in this data directory.
+ *
+ * To prevent two concurrent processes working with the same data
+ * directory, we first try to lock the lockfile exclusively.
+ *
+ * - The file must be stale, probably left over from a previous system
+ * boot cycle. The same if the lockfile contains our parent's or
+ * grandparent's PID.
+ *
+ * We need to check this because of the likelihood that a reboot will
+ * assign exactly the same PID as we had in the previous reboot, or one
+ * that's only one or two counts larger and hence the lockfile's PID now
+ * refers to an ancestor shell process. We allow pg_ctl to pass down its
+ * parent shell PID (our grandparent PID) via the environment variable
+ * PG_GRANDPARENT_PID; this is so that launching the postmaster via pg_ctl
+ * can be just as reliable as launching it directly. There is no
+ * provision for detecting further-removed ancestor processes, but if the
+ * init script is written carefully then all but the immediate parent
+ * shell will be root-owned processes and so the kill test will fail with
+ * EPERM. Note that we cannot get a false negative this way, because an
+ * existing postmaster would surely never launch a competing postmaster or
+ * pg_ctl process directly.
*/
my_pid = getpid();
@@ -1224,7 +1294,11 @@ CreateLockFile(const char *filename, bool amPostmaster,
*/
fd = open(filename, O_RDWR | O_CREAT | O_EXCL, pg_file_create_mode);
if (fd >= 0)
- break; /* Success; exit the retry loop */
+ {
+ /* Success; lock and exit the retry loop */
+ flock_fd = OFDLockFile(fd, filename);
+ break;
+ }
/*
* Couldn't create the pid file. Probably it already exists.
@@ -1238,8 +1312,12 @@ CreateLockFile(const char *filename, bool amPostmaster,
/*
* Read the file to get the old owner's PID. Note race condition
* here: file might have been deleted since we tried to create it.
+ *
+ * We're going to use the same fd for flock, and want to create a
+ * write lock for the latter one. Since both fd and the lock have to
+ * be of the same type, open the file for read and write.
*/
- fd = open(filename, O_RDONLY, pg_file_create_mode);
+ fd = open(filename, O_RDWR, pg_file_create_mode);
if (fd < 0)
{
if (errno == ENOENT)
@@ -1249,6 +1327,10 @@ CreateLockFile(const char *filename, bool amPostmaster,
errmsg("could not open lock file \"%s\": %m",
filename)));
}
+
+ /* Try to lock the file. We stop here, if it's already locked. */
+ flock_fd = OFDLockFile(fd, filename);
+
pgstat_report_wait_start(WAIT_EVENT_LOCK_FILE_CREATE_READ);
if ((len = read(fd, buffer, sizeof(buffer) - 1)) < 0)
ereport(FATAL,
@@ -1448,7 +1530,11 @@ CreateLockFile(const char *filename, bool amPostmaster,
* Use lcons so that the lock files are unlinked in reverse order of
* creation; this is critical!
*/
- lock_files = lcons(pstrdup(filename), lock_files);
+ lock_file = palloc0_object(PGLockFile);
+ lock_file->filename = pstrdup(filename);
+ lock_file->fd = flock_fd;
+
+ lock_files = lcons(lock_file, lock_files);
}
/*
@@ -1495,14 +1581,14 @@ TouchSocketLockFiles(void)
foreach(l, lock_files)
{
- char *socketLockFile = (char *) lfirst(l);
+ PGLockFile *lock_file = (PGLockFile *) lfirst(l);
/* No need to touch the data directory lock file, we trust */
- if (strcmp(socketLockFile, DIRECTORY_LOCK_FILE) == 0)
+ if (strcmp(lock_file->filename, DIRECTORY_LOCK_FILE) == 0)
continue;
/* we just ignore any error here */
- (void) utime(socketLockFile, NULL);
+ (void) utime(lock_file->filename, NULL);
}
}
diff --git a/src/include/pg_config.h.in b/src/include/pg_config.h.in
index 4f8113c144b..cc38c06dc13 100644
--- a/src/include/pg_config.h.in
+++ b/src/include/pg_config.h.in
@@ -85,6 +85,10 @@
don't. */
#undef HAVE_DECL_F_FULLFSYNC
+/* Define to 1 if you have the declaration of `F_OFD_SETLK', and to 0 if you
+ don't. */
+#undef HAVE_DECL_F_OFD_SETLK
+
/* Define to 1 if you have the declaration of `memset_s', and to 0 if you
don't. */
#undef HAVE_DECL_MEMSET_S
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index ffb413ab612..ad34142e7d6 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -1956,6 +1956,7 @@ PGIOAlignedBlock
PGLZ_HistEntry
PGLZ_Strategy
PGLoadBalanceType
+PGLockFile
PGMessageField
PGModuleMagicFunction
PGNoticeHooks
base-commit: 73dfe79fd6034b1e7e41e83d9c82c166dba8eb67
--
2.52.0
Attachments:
[text/plain] v4-0001-Use-open-file-description-locks-for-lockfiles.patch (12.5K, ../../c5jj2zvhnh36xfvqkcogrvk5fkcut3un5q5hibozfm62h6t4kp@sdlorh7uaf2s/2-v4-0001-Use-open-file-description-locks-for-lockfiles.patch)
download | inline diff:
From 2ebad2531437fb1a943e7337981098c1268d087a Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Thu, 18 Dec 2025 18:21:59 +0100
Subject: [PATCH v4] Use open file description locks for lockfiles
When starting up, postmaster checks for an existing data directory lockfile. If
this file contains current process PID, it's assumed to be stale. Turns out
there is another possibility: we might be running in a PID namespace, and there
is another postgres running inside another PID namespace using the same data
directory. The result is that we don't see another process due to namespace
isolation and start concurrently with the other.
To prevent such situations, at startup use fcntl to get an exclusive open file
description lock for data directory lockfile. Since such locks are associated
with open file descriptors, meaning they're not affected by PID namespace
isolation. It's a "best effort" locking, intended to work with already existing
mechanism, not replace it.
This approach was discussed multiple times in the past, and usually was
rejected as the main work horse for the data directory lockfile due to:
* Portability issues. Open file description lock was a non-POSIX extension in
Linux and similar flock is from BSD standard. But looks like everybody agrees
that such locks make more sense than a typical advisory locks, and
F_OFD_SETLK made its way into POSIX.1 2024 [1].
* Issues with NFS. The current state of things here looks like this:
- NFSv3 doesn't implement open file description locks, they're converted to
advisory locks instead. Advisory locks are subject to namespace isolation,
meaning that processes in different PID namespaces will not see each other
advisory lock, and it's still possible to run multiple postgres
instances on the same data directory.
- NFSv4 uses a lease system for locking, I haven't found any mention of
conversion to advisory locks neither in the man page nor in RFC [2].
To summarize, the approach is now considered POSIX and should fix the described
problem everywhere, except NFSv3.
Use open file description lock for both data directory and socker
lockfiles, since both are affected in the same way.
[1]: https://pubs.opengroup.org/onlinepubs/9799919799/functions/fcntl.html
[2]: https://www.rfc-editor.org/rfc/rfc7530
Reviewed-by: Ilmar Yunusov <tanswis42@gmail.com>
---
configure | 14 +++
configure.ac | 3 +
meson.build | 1 +
src/backend/utils/init/miscinit.c | 136 ++++++++++++++++++++++++------
src/include/pg_config.h.in | 4 +
src/tools/pgindent/typedefs.list | 1 +
6 files changed, 134 insertions(+), 25 deletions(-)
diff --git a/configure b/configure
index 35b0b72f0a7..25ebcd3cc47 100755
--- a/configure
+++ b/configure
@@ -16444,6 +16444,20 @@ cat >>confdefs.h <<_ACEOF
_ACEOF
+# Linux open file descriptor locks
+ac_fn_c_check_decl "$LINENO" "F_OFD_SETLK" "ac_cv_have_decl_F_OFD_SETLK" "#include <fcntl.h>
+"
+if test "x$ac_cv_have_decl_F_OFD_SETLK" = xyes; then :
+ ac_have_decl=1
+else
+ ac_have_decl=0
+fi
+
+cat >>confdefs.h <<_ACEOF
+#define HAVE_DECL_F_OFD_SETLK $ac_have_decl
+_ACEOF
+
+
ac_fn_c_check_func "$LINENO" "explicit_bzero" "ac_cv_func_explicit_bzero"
if test "x$ac_cv_func_explicit_bzero" = xyes; then :
$as_echo "#define HAVE_EXPLICIT_BZERO 1" >>confdefs.h
diff --git a/configure.ac b/configure.ac
index 0e624fe36b9..677137207e7 100644
--- a/configure.ac
+++ b/configure.ac
@@ -1913,6 +1913,9 @@ AC_CHECK_DECLS([memset_s], [], [], [#define __STDC_WANT_LIB_EXT1__ 1
# This is probably only present on macOS, but may as well check always
AC_CHECK_DECLS(F_FULLFSYNC, [], [], [#include <fcntl.h>])
+# Linux open file descriptor locks
+AC_CHECK_DECLS([F_OFD_SETLK], [], [], [#include <fcntl.h>])
+
AC_REPLACE_FUNCS(m4_normalize([
explicit_bzero
getopt
diff --git a/meson.build b/meson.build
index d88a7a70308..153fbb477bb 100644
--- a/meson.build
+++ b/meson.build
@@ -2901,6 +2901,7 @@ decl_checks = [
['strlcpy', 'string.h'],
['strsep', 'string.h'],
['timingsafe_bcmp', 'string.h'],
+ ['F_OFD_SETLK', 'fcntl.h'],
]
# Need to check for function declarations for these functions, because
diff --git a/src/backend/utils/init/miscinit.c b/src/backend/utils/init/miscinit.c
index 7ffc808073a..d4c2f80eb46 100644
--- a/src/backend/utils/init/miscinit.c
+++ b/src/backend/utils/init/miscinit.c
@@ -69,6 +69,15 @@ static List *lock_files = NIL;
static Latch LocalLatchData;
+typedef struct
+{
+ /* LockFile name. */
+ const char *filename;
+
+ /* File descriptor for open file description lock. */
+ int fd;
+} PGLockFile;
+
/* ----------------------------------------------------------------
* ignoring system indexes support stuff
*
@@ -1119,6 +1128,48 @@ RestoreClientConnectionInfo(char *conninfo)
*-------------------------------------------------------------------------
*/
+/*
+ * OFD lock the specified lockfile.
+ *
+ * Lock the lockfile with an open file description lock. If the lock is already
+ * taken, it's a hard stop. It's only a best effort test, and any other errors
+ * are ignored. On succes the file descriptor is duplicated, to make sure there
+ * will be at least one open copy of it to keep the lock.
+ *
+ * filename argument is used only for reporting purposes.
+ */
+static int
+OFDLockFile(int fd, const char *filename)
+{
+#if HAVE_DECL_F_OFD_SETLK
+ struct flock lock;
+
+ lock.l_type = F_WRLCK;
+ lock.l_whence = SEEK_SET;
+ lock.l_start = 0;
+ lock.l_len = 0;
+ lock.l_pid = 0;
+
+ if (fcntl(fd, F_OFD_SETLK, &lock) == -1)
+ {
+ if (errno == EAGAIN)
+ ereport(FATAL,
+ (errcode(ERRCODE_LOCK_FILE_EXISTS),
+ errmsg("cannot lock the lock file \"%s\"", filename),
+ errhint("Another server is starting.")));
+ else
+ {
+ elog(WARNING, "Failed locking file \"%s\", %m", filename);
+ return -1;
+ }
+ }
+ else
+ return dup(fd);
+#else
+ return -1;
+#endif
+}
+
/*
* proc_exit callback to remove lockfiles.
*/
@@ -1129,9 +1180,16 @@ UnlinkLockFiles(int status, Datum arg)
foreach(l, lock_files)
{
- char *curfile = (char *) lfirst(l);
+ PGLockFile *lock_file = (PGLockFile *) lfirst(l);
- unlink(curfile);
+ /*
+ * Close the file descriptor, which keeps the open file description
+ * lock.
+ */
+ if (lock_file->fd > 0)
+ close(lock_file->fd);
+
+ unlink(lock_file->filename);
/* Should we complain if the unlink fails? */
}
/* Since we're about to exit, no need to reclaim storage */
@@ -1161,7 +1219,9 @@ CreateLockFile(const char *filename, bool amPostmaster,
const char *socketDir,
bool isDDLock, const char *refName)
{
- int fd;
+ int fd,
+ flock_fd = -1;
+ PGLockFile *lock_file;
char buffer[MAXPGPATH * 2 + 256];
int ntries;
int len;
@@ -1173,22 +1233,32 @@ CreateLockFile(const char *filename, bool amPostmaster,
const char *envvar;
/*
- * If the PID in the lockfile is our own PID or our parent's or
- * grandparent's PID, then the file must be stale (probably left over from
- * a previous system boot cycle). We need to check this because of the
- * likelihood that a reboot will assign exactly the same PID as we had in
- * the previous reboot, or one that's only one or two counts larger and
- * hence the lockfile's PID now refers to an ancestor shell process. We
- * allow pg_ctl to pass down its parent shell PID (our grandparent PID)
- * via the environment variable PG_GRANDPARENT_PID; this is so that
- * launching the postmaster via pg_ctl can be just as reliable as
- * launching it directly. There is no provision for detecting
- * further-removed ancestor processes, but if the init script is written
- * carefully then all but the immediate parent shell will be root-owned
- * processes and so the kill test will fail with EPERM. Note that we
- * cannot get a false negative this way, because an existing postmaster
- * would surely never launch a competing postmaster or pg_ctl process
- * directly.
+ * If we find an already existing lockfile containing our own PID, there
+ * are few options:
+ *
+ * - There is another process, that we don't see due to PID namespace
+ * isolation, which is already running in this data directory.
+ *
+ * To prevent two concurrent processes working with the same data
+ * directory, we first try to lock the lockfile exclusively.
+ *
+ * - The file must be stale, probably left over from a previous system
+ * boot cycle. The same if the lockfile contains our parent's or
+ * grandparent's PID.
+ *
+ * We need to check this because of the likelihood that a reboot will
+ * assign exactly the same PID as we had in the previous reboot, or one
+ * that's only one or two counts larger and hence the lockfile's PID now
+ * refers to an ancestor shell process. We allow pg_ctl to pass down its
+ * parent shell PID (our grandparent PID) via the environment variable
+ * PG_GRANDPARENT_PID; this is so that launching the postmaster via pg_ctl
+ * can be just as reliable as launching it directly. There is no
+ * provision for detecting further-removed ancestor processes, but if the
+ * init script is written carefully then all but the immediate parent
+ * shell will be root-owned processes and so the kill test will fail with
+ * EPERM. Note that we cannot get a false negative this way, because an
+ * existing postmaster would surely never launch a competing postmaster or
+ * pg_ctl process directly.
*/
my_pid = getpid();
@@ -1224,7 +1294,11 @@ CreateLockFile(const char *filename, bool amPostmaster,
*/
fd = open(filename, O_RDWR | O_CREAT | O_EXCL, pg_file_create_mode);
if (fd >= 0)
- break; /* Success; exit the retry loop */
+ {
+ /* Success; lock and exit the retry loop */
+ flock_fd = OFDLockFile(fd, filename);
+ break;
+ }
/*
* Couldn't create the pid file. Probably it already exists.
@@ -1238,8 +1312,12 @@ CreateLockFile(const char *filename, bool amPostmaster,
/*
* Read the file to get the old owner's PID. Note race condition
* here: file might have been deleted since we tried to create it.
+ *
+ * We're going to use the same fd for flock, and want to create a
+ * write lock for the latter one. Since both fd and the lock have to
+ * be of the same type, open the file for read and write.
*/
- fd = open(filename, O_RDONLY, pg_file_create_mode);
+ fd = open(filename, O_RDWR, pg_file_create_mode);
if (fd < 0)
{
if (errno == ENOENT)
@@ -1249,6 +1327,10 @@ CreateLockFile(const char *filename, bool amPostmaster,
errmsg("could not open lock file \"%s\": %m",
filename)));
}
+
+ /* Try to lock the file. We stop here, if it's already locked. */
+ flock_fd = OFDLockFile(fd, filename);
+
pgstat_report_wait_start(WAIT_EVENT_LOCK_FILE_CREATE_READ);
if ((len = read(fd, buffer, sizeof(buffer) - 1)) < 0)
ereport(FATAL,
@@ -1448,7 +1530,11 @@ CreateLockFile(const char *filename, bool amPostmaster,
* Use lcons so that the lock files are unlinked in reverse order of
* creation; this is critical!
*/
- lock_files = lcons(pstrdup(filename), lock_files);
+ lock_file = palloc0_object(PGLockFile);
+ lock_file->filename = pstrdup(filename);
+ lock_file->fd = flock_fd;
+
+ lock_files = lcons(lock_file, lock_files);
}
/*
@@ -1495,14 +1581,14 @@ TouchSocketLockFiles(void)
foreach(l, lock_files)
{
- char *socketLockFile = (char *) lfirst(l);
+ PGLockFile *lock_file = (PGLockFile *) lfirst(l);
/* No need to touch the data directory lock file, we trust */
- if (strcmp(socketLockFile, DIRECTORY_LOCK_FILE) == 0)
+ if (strcmp(lock_file->filename, DIRECTORY_LOCK_FILE) == 0)
continue;
/* we just ignore any error here */
- (void) utime(socketLockFile, NULL);
+ (void) utime(lock_file->filename, NULL);
}
}
diff --git a/src/include/pg_config.h.in b/src/include/pg_config.h.in
index 4f8113c144b..cc38c06dc13 100644
--- a/src/include/pg_config.h.in
+++ b/src/include/pg_config.h.in
@@ -85,6 +85,10 @@
don't. */
#undef HAVE_DECL_F_FULLFSYNC
+/* Define to 1 if you have the declaration of `F_OFD_SETLK', and to 0 if you
+ don't. */
+#undef HAVE_DECL_F_OFD_SETLK
+
/* Define to 1 if you have the declaration of `memset_s', and to 0 if you
don't. */
#undef HAVE_DECL_MEMSET_S
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index ffb413ab612..ad34142e7d6 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -1956,6 +1956,7 @@ PGIOAlignedBlock
PGLZ_HistEntry
PGLZ_Strategy
PGLoadBalanceType
+PGLockFile
PGMessageField
PGModuleMagicFunction
PGNoticeHooks
base-commit: 73dfe79fd6034b1e7e41e83d9c82c166dba8eb67
--
2.52.0
^ permalink raw reply [nested|flat] 10+ messages in thread
* Re: File locks for data directory lockfile in the context of Linux namespaces
2025-12-19 14:27 File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
2026-01-17 15:26 ` Re: File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
2026-06-05 12:37 ` Re: File locks for data directory lockfile in the context of Linux namespaces Ilmar Yunusov <tanswis42@gmail.com>
2026-06-19 15:11 ` Re: File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
2026-06-23 14:29 ` Re: File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
2026-07-06 07:30 ` Re: File locks for data directory lockfile in the context of Linux namespaces Ilmar Yunusov <tanswis42@gmail.com>
2026-07-06 12:10 ` Re: File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
@ 2026-07-06 22:07 ` Zsolt Parragi <zsolt.parragi@percona.com>
2026-07-10 10:20 ` Re: File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 1 reply; 10+ messages in thread
From: Zsolt Parragi @ 2026-07-06 22:07 UTC (permalink / raw)
To: pgsql-hackers@lists.postgresql.org
Hello!
+ if (errno == EAGAIN)
+ ereport(FATAL,
+ (errcode(ERRCODE_LOCK_FILE_EXISTS),
According to fcntl.2, this should handle both EACCESS and EAGAIN:
ERRORS
EACCES or EAGAIN
Operation is prohibited by locks held by other processes.
+static int
+OFDLockFile(int fd, const char *filename)
+...
+ else
+ return dup(fd);
Isn't this missing an FD_CLOEXEC, so that launched processes doesn't
inherit it and keep the lock open possibly longer than needed?
Also, shouldn't the code verify the result of dup? (!= -1 / errno)
+ flock_fd = OFDLockFile(fd, filename);
Can't we leak flock_fd in the stale path?
+ * Close the file descriptor, which keeps the open file description
+ * lock.
+ */
+ if (lock_file->fd > 0)
+ close(lock_file->fd);
Shouldn't this check for >= 0?
+ elog(WARNING, "Failed locking file \"%s\", %m", filename);
This probably should be:
ereport(WARNING, (errcode_for_file_access(), errmsg("could not lock
file \"%s\": %m", filename)))
^ permalink raw reply [nested|flat] 10+ messages in thread
* Re: File locks for data directory lockfile in the context of Linux namespaces
2025-12-19 14:27 File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
2026-01-17 15:26 ` Re: File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
2026-06-05 12:37 ` Re: File locks for data directory lockfile in the context of Linux namespaces Ilmar Yunusov <tanswis42@gmail.com>
2026-06-19 15:11 ` Re: File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
2026-06-23 14:29 ` Re: File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
2026-07-06 07:30 ` Re: File locks for data directory lockfile in the context of Linux namespaces Ilmar Yunusov <tanswis42@gmail.com>
2026-07-06 12:10 ` Re: File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
2026-07-06 22:07 ` Re: File locks for data directory lockfile in the context of Linux namespaces Zsolt Parragi <zsolt.parragi@percona.com>
@ 2026-07-10 10:20 ` Dmitry Dolgov <9erthalion6@gmail.com>
2026-08-28 07:28 ` Re: File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 1 reply; 10+ messages in thread
From: Dmitry Dolgov @ 2026-07-10 10:20 UTC (permalink / raw)
To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: pgsql-hackers@lists.postgresql.org
> On Mon, Jul 06, 2026 at 03:07:35PM -0700, Zsolt Parragi wrote:
> + if (errno == EAGAIN)
> + ereport(FATAL,
> + (errcode(ERRCODE_LOCK_FILE_EXISTS),
>
> According to fcntl.2, this should handle both EACCESS and EAGAIN:
>
> ERRORS
> EACCES or EAGAIN
> Operation is prohibited by locks held by other processes.
Good point. From what I see after a cursory look at fcntl is that it
normally returns EAGAIN, but filesystems are allowed to implement a
custom lock operation, so it makes sense to be prepared.
> +static int
> +OFDLockFile(int fd, const char *filename)
> +...
> + else
> + return dup(fd);
>
> Isn't this missing an FD_CLOEXEC, so that launched processes doesn't
> inherit it and keep the lock open possibly longer than needed?
This is an interesting question. I haven't thought about this
originally, but now I think the current approach (no FD_CLOEXEC) is what
is actually needed. We want to keep the lock as long as any existing
process may access the data directory, thus the lock lifetime must be
equal to the lifetime of a longest living process. Currently the lock
file is created by the bootstrap process and the postmaster, which I
think fits the picture.
> Also, shouldn't the code verify the result of dup? (!= -1 / errno)
From the functionality perspective it's not needed, but yeah, it will
make reporting better.
>
> + flock_fd = OFDLockFile(fd, filename);
>
> Can't we leak flock_fd in the stale path?
How, do you see any particular scenario?
> + elog(WARNING, "Failed locking file \"%s\", %m", filename);
>
> This probably should be:
>
> ereport(WARNING, (errcode_for_file_access(), errmsg("could not lock
> file \"%s\": %m", filename)))
I would concider this one of those "shouldn't be possible" errors, and
as such elog is more appropriate here.
^ permalink raw reply [nested|flat] 10+ messages in thread
* Re: File locks for data directory lockfile in the context of Linux namespaces
2025-12-19 14:27 File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
2026-01-17 15:26 ` Re: File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
2026-06-05 12:37 ` Re: File locks for data directory lockfile in the context of Linux namespaces Ilmar Yunusov <tanswis42@gmail.com>
2026-06-19 15:11 ` Re: File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
2026-06-23 14:29 ` Re: File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
2026-07-06 07:30 ` Re: File locks for data directory lockfile in the context of Linux namespaces Ilmar Yunusov <tanswis42@gmail.com>
2026-07-06 12:10 ` Re: File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
2026-07-06 22:07 ` Re: File locks for data directory lockfile in the context of Linux namespaces Zsolt Parragi <zsolt.parragi@percona.com>
2026-07-10 10:20 ` Re: File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
@ 2026-08-28 07:28 ` Dmitry Dolgov <9erthalion6@gmail.com>
0 siblings, 0 replies; 10+ messages in thread
From: Dmitry Dolgov @ 2026-08-28 07:28 UTC (permalink / raw)
To: Zsolt Parragi <zsolt.parragi@percona.com>; +Cc: pgsql-hackers@lists.postgresql.org
Before I completely forgot about this one, here are the discussed
changes.
From 3b7305030be559fcdbf1f03cedd599bfa00c0574 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Thu, 18 Dec 2025 18:21:59 +0100
Subject: [PATCH v5] Use open file description locks for lockfiles
When starting up, postmaster checks for an existing data directory lockfile. If
this file contains current process PID, it's assumed to be stale. Turns out
there is another possibility: we might be running in a PID namespace, and there
is another postgres running inside another PID namespace using the same data
directory. The result is that we don't see another process due to namespace
isolation and start concurrently with the other.
To prevent such situations, at startup use fcntl to get an exclusive open file
description lock for data directory lockfile. Since such locks are associated
with open file descriptors, meaning they're not affected by PID namespace
isolation. It's a "best effort" locking, intended to work with already existing
mechanism, not replace it.
This approach was discussed multiple times in the past, and usually was
rejected as the main work horse for the data directory lockfile due to:
* Portability issues. Open file description lock was a non-POSIX extension in
Linux and similar flock is from BSD standard. But looks like everybody agrees
that such locks make more sense than a typical advisory locks, and
F_OFD_SETLK made its way into POSIX.1 2024 [1].
* Issues with NFS. The current state of things here looks like this:
- NFSv3 doesn't implement open file description locks, they're converted to
advisory locks instead. Advisory locks are subject to namespace isolation,
meaning that processes in different PID namespaces will not see each other
advisory lock, and it's still possible to run multiple postgres
instances on the same data directory.
- NFSv4 uses a lease system for locking, I haven't found any mention of
conversion to advisory locks neither in the man page nor in RFC [2].
To summarize, the approach is now considered POSIX and should fix the described
problem everywhere, except NFSv3.
Use open file description lock for both data directory and socker
lockfiles, since both are affected in the same way. Note that we
intentionally do not use FD_CLOEXEC to inherit the lock fd, since this
way the lock is kept for as long as it's needed.
[1]: https://pubs.opengroup.org/onlinepubs/9799919799/functions/fcntl.html
[2]: https://www.rfc-editor.org/rfc/rfc7530
Reviewed-by: Ilmar Yunusov <tanswis42@gmail.com>
Reviewed-by: Zsolt Parragi <zsolt.parragi@percona.com>
---
configure | 14 +++
configure.ac | 3 +
meson.build | 1 +
src/backend/utils/init/miscinit.c | 150 +++++++++++++++++++++++++-----
src/include/pg_config.h.in | 4 +
src/tools/pgindent/typedefs.list | 1 +
6 files changed, 148 insertions(+), 25 deletions(-)
diff --git a/configure b/configure
index d42a7a794ff..a072b628855 100755
--- a/configure
+++ b/configure
@@ -16414,6 +16414,20 @@ cat >>confdefs.h <<_ACEOF
_ACEOF
+# Linux open file descriptor locks
+ac_fn_c_check_decl "$LINENO" "F_OFD_SETLK" "ac_cv_have_decl_F_OFD_SETLK" "#include <fcntl.h>
+"
+if test "x$ac_cv_have_decl_F_OFD_SETLK" = xyes; then :
+ ac_have_decl=1
+else
+ ac_have_decl=0
+fi
+
+cat >>confdefs.h <<_ACEOF
+#define HAVE_DECL_F_OFD_SETLK $ac_have_decl
+_ACEOF
+
+
ac_fn_c_check_func "$LINENO" "explicit_bzero" "ac_cv_func_explicit_bzero"
if test "x$ac_cv_func_explicit_bzero" = xyes; then :
$as_echo "#define HAVE_EXPLICIT_BZERO 1" >>confdefs.h
diff --git a/configure.ac b/configure.ac
index a331749fcb5..29fe4a077c4 100644
--- a/configure.ac
+++ b/configure.ac
@@ -1912,6 +1912,9 @@ AC_CHECK_DECLS([memset_s], [], [], [#define __STDC_WANT_LIB_EXT1__ 1
# This is probably only present on macOS, but may as well check always
AC_CHECK_DECLS(F_FULLFSYNC, [], [], [#include <fcntl.h>])
+# Linux open file descriptor locks
+AC_CHECK_DECLS([F_OFD_SETLK], [], [], [#include <fcntl.h>])
+
AC_REPLACE_FUNCS(m4_normalize([
explicit_bzero
getopt
diff --git a/meson.build b/meson.build
index f4cde249242..8edd8cda474 100644
--- a/meson.build
+++ b/meson.build
@@ -2892,6 +2892,7 @@ decl_checks = [
['strlcpy', 'string.h'],
['strsep', 'string.h'],
['timingsafe_bcmp', 'string.h'],
+ ['F_OFD_SETLK', 'fcntl.h'],
]
# Need to check for function declarations for these functions, because
diff --git a/src/backend/utils/init/miscinit.c b/src/backend/utils/init/miscinit.c
index eddce1ce33f..334031aed0a 100644
--- a/src/backend/utils/init/miscinit.c
+++ b/src/backend/utils/init/miscinit.c
@@ -69,6 +69,15 @@ static List *lock_files = NIL;
static Latch LocalLatchData;
+typedef struct
+{
+ /* LockFile name. */
+ const char *filename;
+
+ /* File descriptor for open file description lock. */
+ int fd;
+} PGLockFile;
+
/* ----------------------------------------------------------------
* ignoring system indexes support stuff
*
@@ -1119,6 +1128,62 @@ RestoreClientConnectionInfo(char *conninfo)
*-------------------------------------------------------------------------
*/
+/*
+ * OFD lock the specified lockfile.
+ *
+ * Lock the lockfile with an open file description lock. If the lock is already
+ * taken, it's a hard stop. It's only a best effort test, and any other errors
+ * are ignored. On succes the file descriptor is duplicated, to make sure there
+ * will be at least one open copy of it to keep the lock.
+ *
+ * filename argument is used only for reporting purposes.
+ */
+static int
+OFDLockFile(int fd, const char *filename)
+{
+#if HAVE_DECL_F_OFD_SETLK
+ struct flock lock;
+
+ lock.l_type = F_WRLCK;
+ lock.l_whence = SEEK_SET;
+ lock.l_start = 0;
+ lock.l_len = 0;
+ lock.l_pid = 0;
+
+ if (fcntl(fd, F_OFD_SETLK, &lock) == -1)
+ {
+ if (errno == EAGAIN || errno == EACCES)
+ ereport(FATAL,
+ (errcode(ERRCODE_LOCK_FILE_EXISTS),
+ errmsg("cannot lock the lock file \"%s\"", filename),
+ errhint("Another server is starting.")));
+ else
+ {
+ elog(WARNING, "Failed locking file \"%s\", %m", filename);
+ return -1;
+ }
+ }
+ else
+ {
+ int dup_fd = dup(fd);
+
+ if (dup_fd < 0)
+ elog(WARNING, "Failed duplicating fd for the lock file \"%s\", %m",
+ filename);
+
+ /*
+ * Note that we intentionally do not set FD_CLOEXEC flag on this fd.
+ * It's lifetime represents lifetime of the corresponding OFD lock, and
+ * the idea is to keep the lock for as long as a longest living
+ * process, including any subprograms.
+ */
+ return dup_fd;
+ }
+#else
+ return -1;
+#endif
+}
+
/*
* proc_exit callback to remove lockfiles.
*/
@@ -1129,9 +1194,16 @@ UnlinkLockFiles(int status, Datum arg)
foreach(l, lock_files)
{
- char *curfile = (char *) lfirst(l);
+ PGLockFile *lock_file = (PGLockFile *) lfirst(l);
- unlink(curfile);
+ /*
+ * Close the file descriptor, which keeps the open file description
+ * lock.
+ */
+ if (lock_file->fd >= 0)
+ close(lock_file->fd);
+
+ unlink(lock_file->filename);
/* Should we complain if the unlink fails? */
}
/* Since we're about to exit, no need to reclaim storage */
@@ -1161,7 +1233,9 @@ CreateLockFile(const char *filename, bool amPostmaster,
const char *socketDir,
bool isDDLock, const char *refName)
{
- int fd;
+ int fd,
+ flock_fd = -1;
+ PGLockFile *lock_file;
char buffer[MAXPGPATH * 2 + 256];
int ntries;
int encoded_pid;
@@ -1172,22 +1246,32 @@ CreateLockFile(const char *filename, bool amPostmaster,
const char *envvar;
/*
- * If the PID in the lockfile is our own PID or our parent's or
- * grandparent's PID, then the file must be stale (probably left over from
- * a previous system boot cycle). We need to check this because of the
- * likelihood that a reboot will assign exactly the same PID as we had in
- * the previous reboot, or one that's only one or two counts larger and
- * hence the lockfile's PID now refers to an ancestor shell process. We
- * allow pg_ctl to pass down its parent shell PID (our grandparent PID)
- * via the environment variable PG_GRANDPARENT_PID; this is so that
- * launching the postmaster via pg_ctl can be just as reliable as
- * launching it directly. There is no provision for detecting
- * further-removed ancestor processes, but if the init script is written
- * carefully then all but the immediate parent shell will be root-owned
- * processes and so the kill test will fail with EPERM. Note that we
- * cannot get a false negative this way, because an existing postmaster
- * would surely never launch a competing postmaster or pg_ctl process
- * directly.
+ * If we find an already existing lockfile containing our own PID, there
+ * are few options:
+ *
+ * - There is another process, that we don't see due to PID namespace
+ * isolation, which is already running in this data directory.
+ *
+ * To prevent two concurrent processes working with the same data
+ * directory, we first try to lock the lockfile exclusively.
+ *
+ * - The file must be stale, probably left over from a previous system
+ * boot cycle. The same if the lockfile contains our parent's or
+ * grandparent's PID.
+ *
+ * We need to check this because of the likelihood that a reboot will
+ * assign exactly the same PID as we had in the previous reboot, or one
+ * that's only one or two counts larger and hence the lockfile's PID now
+ * refers to an ancestor shell process. We allow pg_ctl to pass down its
+ * parent shell PID (our grandparent PID) via the environment variable
+ * PG_GRANDPARENT_PID; this is so that launching the postmaster via pg_ctl
+ * can be just as reliable as launching it directly. There is no
+ * provision for detecting further-removed ancestor processes, but if the
+ * init script is written carefully then all but the immediate parent
+ * shell will be root-owned processes and so the kill test will fail with
+ * EPERM. Note that we cannot get a false negative this way, because an
+ * existing postmaster would surely never launch a competing postmaster or
+ * pg_ctl process directly.
*/
my_pid = getpid();
@@ -1225,7 +1309,11 @@ CreateLockFile(const char *filename, bool amPostmaster,
*/
fd = open(filename, O_RDWR | O_CREAT | O_EXCL, pg_file_create_mode);
if (fd >= 0)
- break; /* Success; exit the retry loop */
+ {
+ /* Success; lock and exit the retry loop */
+ flock_fd = OFDLockFile(fd, filename);
+ break;
+ }
/*
* Couldn't create the pid file. Probably it already exists.
@@ -1239,8 +1327,12 @@ CreateLockFile(const char *filename, bool amPostmaster,
/*
* Read the file to get the old owner's PID. Note race condition
* here: file might have been deleted since we tried to create it.
+ *
+ * We're going to use the same fd for flock, and want to create a
+ * write lock for the latter one. Since both fd and the lock have to
+ * be of the same type, open the file for read and write.
*/
- fd = open(filename, O_RDONLY, pg_file_create_mode);
+ fd = open(filename, O_RDWR, pg_file_create_mode);
if (fd < 0)
{
if (errno == ENOENT)
@@ -1250,6 +1342,10 @@ CreateLockFile(const char *filename, bool amPostmaster,
errmsg("could not open lock file \"%s\": %m",
filename)));
}
+
+ /* Try to lock the file. We stop here, if it's already locked. */
+ flock_fd = OFDLockFile(fd, filename);
+
pgstat_report_wait_start(WAIT_EVENT_LOCK_FILE_CREATE_READ);
if ((len = read(fd, buffer, sizeof(buffer) - 1)) < 0)
ereport(FATAL,
@@ -1449,7 +1545,11 @@ CreateLockFile(const char *filename, bool amPostmaster,
* Use lcons so that the lock files are unlinked in reverse order of
* creation; this is critical!
*/
- lock_files = lcons(pstrdup(filename), lock_files);
+ lock_file = palloc0_object(PGLockFile);
+ lock_file->filename = pstrdup(filename);
+ lock_file->fd = flock_fd;
+
+ lock_files = lcons(lock_file, lock_files);
}
/*
@@ -1496,14 +1596,14 @@ TouchSocketLockFiles(void)
foreach(l, lock_files)
{
- char *socketLockFile = (char *) lfirst(l);
+ PGLockFile *lock_file = (PGLockFile *) lfirst(l);
/* No need to touch the data directory lock file, we trust */
- if (strcmp(socketLockFile, DIRECTORY_LOCK_FILE) == 0)
+ if (strcmp(lock_file->filename, DIRECTORY_LOCK_FILE) == 0)
continue;
/* we just ignore any error here */
- (void) utime(socketLockFile, NULL);
+ (void) utime(lock_file->filename, NULL);
}
}
diff --git a/src/include/pg_config.h.in b/src/include/pg_config.h.in
index 661c4a9b168..fae00af3284 100644
--- a/src/include/pg_config.h.in
+++ b/src/include/pg_config.h.in
@@ -85,6 +85,10 @@
don't. */
#undef HAVE_DECL_F_FULLFSYNC
+/* Define to 1 if you have the declaration of `F_OFD_SETLK', and to 0 if you
+ don't. */
+#undef HAVE_DECL_F_OFD_SETLK
+
/* Define to 1 if you have the declaration of `memset_s', and to 0 if you
don't. */
#undef HAVE_DECL_MEMSET_S
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 59cf40b5bb0..fdd2ff2cc32 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -1958,6 +1958,7 @@ PGIOAlignedBlock
PGLZ_HistEntry
PGLZ_Strategy
PGLoadBalanceType
+PGLockFile
PGMessageField
PGModuleMagicFunction
PGNoticeHooks
base-commit: 45817676632b46fc9d2b97c8d432348e0ee4f1f4
--
2.55.0
Attachments:
[text/plain] v5-0001-Use-open-file-description-locks-for-lockfiles.patch (13.1K, ../../apE2v2ygGkuAr56a@ddolgov-thinkpadt14sgen1.rmtde.csb/2-v5-0001-Use-open-file-description-locks-for-lockfiles.patch)
download | inline diff:
From 3b7305030be559fcdbf1f03cedd599bfa00c0574 Mon Sep 17 00:00:00 2001
From: Dmitrii Dolgov <9erthalion6@gmail.com>
Date: Thu, 18 Dec 2025 18:21:59 +0100
Subject: [PATCH v5] Use open file description locks for lockfiles
When starting up, postmaster checks for an existing data directory lockfile. If
this file contains current process PID, it's assumed to be stale. Turns out
there is another possibility: we might be running in a PID namespace, and there
is another postgres running inside another PID namespace using the same data
directory. The result is that we don't see another process due to namespace
isolation and start concurrently with the other.
To prevent such situations, at startup use fcntl to get an exclusive open file
description lock for data directory lockfile. Since such locks are associated
with open file descriptors, meaning they're not affected by PID namespace
isolation. It's a "best effort" locking, intended to work with already existing
mechanism, not replace it.
This approach was discussed multiple times in the past, and usually was
rejected as the main work horse for the data directory lockfile due to:
* Portability issues. Open file description lock was a non-POSIX extension in
Linux and similar flock is from BSD standard. But looks like everybody agrees
that such locks make more sense than a typical advisory locks, and
F_OFD_SETLK made its way into POSIX.1 2024 [1].
* Issues with NFS. The current state of things here looks like this:
- NFSv3 doesn't implement open file description locks, they're converted to
advisory locks instead. Advisory locks are subject to namespace isolation,
meaning that processes in different PID namespaces will not see each other
advisory lock, and it's still possible to run multiple postgres
instances on the same data directory.
- NFSv4 uses a lease system for locking, I haven't found any mention of
conversion to advisory locks neither in the man page nor in RFC [2].
To summarize, the approach is now considered POSIX and should fix the described
problem everywhere, except NFSv3.
Use open file description lock for both data directory and socker
lockfiles, since both are affected in the same way. Note that we
intentionally do not use FD_CLOEXEC to inherit the lock fd, since this
way the lock is kept for as long as it's needed.
[1]: https://pubs.opengroup.org/onlinepubs/9799919799/functions/fcntl.html
[2]: https://www.rfc-editor.org/rfc/rfc7530
Reviewed-by: Ilmar Yunusov <tanswis42@gmail.com>
Reviewed-by: Zsolt Parragi <zsolt.parragi@percona.com>
---
configure | 14 +++
configure.ac | 3 +
meson.build | 1 +
src/backend/utils/init/miscinit.c | 150 +++++++++++++++++++++++++-----
src/include/pg_config.h.in | 4 +
src/tools/pgindent/typedefs.list | 1 +
6 files changed, 148 insertions(+), 25 deletions(-)
diff --git a/configure b/configure
index d42a7a794ff..a072b628855 100755
--- a/configure
+++ b/configure
@@ -16414,6 +16414,20 @@ cat >>confdefs.h <<_ACEOF
_ACEOF
+# Linux open file descriptor locks
+ac_fn_c_check_decl "$LINENO" "F_OFD_SETLK" "ac_cv_have_decl_F_OFD_SETLK" "#include <fcntl.h>
+"
+if test "x$ac_cv_have_decl_F_OFD_SETLK" = xyes; then :
+ ac_have_decl=1
+else
+ ac_have_decl=0
+fi
+
+cat >>confdefs.h <<_ACEOF
+#define HAVE_DECL_F_OFD_SETLK $ac_have_decl
+_ACEOF
+
+
ac_fn_c_check_func "$LINENO" "explicit_bzero" "ac_cv_func_explicit_bzero"
if test "x$ac_cv_func_explicit_bzero" = xyes; then :
$as_echo "#define HAVE_EXPLICIT_BZERO 1" >>confdefs.h
diff --git a/configure.ac b/configure.ac
index a331749fcb5..29fe4a077c4 100644
--- a/configure.ac
+++ b/configure.ac
@@ -1912,6 +1912,9 @@ AC_CHECK_DECLS([memset_s], [], [], [#define __STDC_WANT_LIB_EXT1__ 1
# This is probably only present on macOS, but may as well check always
AC_CHECK_DECLS(F_FULLFSYNC, [], [], [#include <fcntl.h>])
+# Linux open file descriptor locks
+AC_CHECK_DECLS([F_OFD_SETLK], [], [], [#include <fcntl.h>])
+
AC_REPLACE_FUNCS(m4_normalize([
explicit_bzero
getopt
diff --git a/meson.build b/meson.build
index f4cde249242..8edd8cda474 100644
--- a/meson.build
+++ b/meson.build
@@ -2892,6 +2892,7 @@ decl_checks = [
['strlcpy', 'string.h'],
['strsep', 'string.h'],
['timingsafe_bcmp', 'string.h'],
+ ['F_OFD_SETLK', 'fcntl.h'],
]
# Need to check for function declarations for these functions, because
diff --git a/src/backend/utils/init/miscinit.c b/src/backend/utils/init/miscinit.c
index eddce1ce33f..334031aed0a 100644
--- a/src/backend/utils/init/miscinit.c
+++ b/src/backend/utils/init/miscinit.c
@@ -69,6 +69,15 @@ static List *lock_files = NIL;
static Latch LocalLatchData;
+typedef struct
+{
+ /* LockFile name. */
+ const char *filename;
+
+ /* File descriptor for open file description lock. */
+ int fd;
+} PGLockFile;
+
/* ----------------------------------------------------------------
* ignoring system indexes support stuff
*
@@ -1119,6 +1128,62 @@ RestoreClientConnectionInfo(char *conninfo)
*-------------------------------------------------------------------------
*/
+/*
+ * OFD lock the specified lockfile.
+ *
+ * Lock the lockfile with an open file description lock. If the lock is already
+ * taken, it's a hard stop. It's only a best effort test, and any other errors
+ * are ignored. On succes the file descriptor is duplicated, to make sure there
+ * will be at least one open copy of it to keep the lock.
+ *
+ * filename argument is used only for reporting purposes.
+ */
+static int
+OFDLockFile(int fd, const char *filename)
+{
+#if HAVE_DECL_F_OFD_SETLK
+ struct flock lock;
+
+ lock.l_type = F_WRLCK;
+ lock.l_whence = SEEK_SET;
+ lock.l_start = 0;
+ lock.l_len = 0;
+ lock.l_pid = 0;
+
+ if (fcntl(fd, F_OFD_SETLK, &lock) == -1)
+ {
+ if (errno == EAGAIN || errno == EACCES)
+ ereport(FATAL,
+ (errcode(ERRCODE_LOCK_FILE_EXISTS),
+ errmsg("cannot lock the lock file \"%s\"", filename),
+ errhint("Another server is starting.")));
+ else
+ {
+ elog(WARNING, "Failed locking file \"%s\", %m", filename);
+ return -1;
+ }
+ }
+ else
+ {
+ int dup_fd = dup(fd);
+
+ if (dup_fd < 0)
+ elog(WARNING, "Failed duplicating fd for the lock file \"%s\", %m",
+ filename);
+
+ /*
+ * Note that we intentionally do not set FD_CLOEXEC flag on this fd.
+ * It's lifetime represents lifetime of the corresponding OFD lock, and
+ * the idea is to keep the lock for as long as a longest living
+ * process, including any subprograms.
+ */
+ return dup_fd;
+ }
+#else
+ return -1;
+#endif
+}
+
/*
* proc_exit callback to remove lockfiles.
*/
@@ -1129,9 +1194,16 @@ UnlinkLockFiles(int status, Datum arg)
foreach(l, lock_files)
{
- char *curfile = (char *) lfirst(l);
+ PGLockFile *lock_file = (PGLockFile *) lfirst(l);
- unlink(curfile);
+ /*
+ * Close the file descriptor, which keeps the open file description
+ * lock.
+ */
+ if (lock_file->fd >= 0)
+ close(lock_file->fd);
+
+ unlink(lock_file->filename);
/* Should we complain if the unlink fails? */
}
/* Since we're about to exit, no need to reclaim storage */
@@ -1161,7 +1233,9 @@ CreateLockFile(const char *filename, bool amPostmaster,
const char *socketDir,
bool isDDLock, const char *refName)
{
- int fd;
+ int fd,
+ flock_fd = -1;
+ PGLockFile *lock_file;
char buffer[MAXPGPATH * 2 + 256];
int ntries;
int encoded_pid;
@@ -1172,22 +1246,32 @@ CreateLockFile(const char *filename, bool amPostmaster,
const char *envvar;
/*
- * If the PID in the lockfile is our own PID or our parent's or
- * grandparent's PID, then the file must be stale (probably left over from
- * a previous system boot cycle). We need to check this because of the
- * likelihood that a reboot will assign exactly the same PID as we had in
- * the previous reboot, or one that's only one or two counts larger and
- * hence the lockfile's PID now refers to an ancestor shell process. We
- * allow pg_ctl to pass down its parent shell PID (our grandparent PID)
- * via the environment variable PG_GRANDPARENT_PID; this is so that
- * launching the postmaster via pg_ctl can be just as reliable as
- * launching it directly. There is no provision for detecting
- * further-removed ancestor processes, but if the init script is written
- * carefully then all but the immediate parent shell will be root-owned
- * processes and so the kill test will fail with EPERM. Note that we
- * cannot get a false negative this way, because an existing postmaster
- * would surely never launch a competing postmaster or pg_ctl process
- * directly.
+ * If we find an already existing lockfile containing our own PID, there
+ * are few options:
+ *
+ * - There is another process, that we don't see due to PID namespace
+ * isolation, which is already running in this data directory.
+ *
+ * To prevent two concurrent processes working with the same data
+ * directory, we first try to lock the lockfile exclusively.
+ *
+ * - The file must be stale, probably left over from a previous system
+ * boot cycle. The same if the lockfile contains our parent's or
+ * grandparent's PID.
+ *
+ * We need to check this because of the likelihood that a reboot will
+ * assign exactly the same PID as we had in the previous reboot, or one
+ * that's only one or two counts larger and hence the lockfile's PID now
+ * refers to an ancestor shell process. We allow pg_ctl to pass down its
+ * parent shell PID (our grandparent PID) via the environment variable
+ * PG_GRANDPARENT_PID; this is so that launching the postmaster via pg_ctl
+ * can be just as reliable as launching it directly. There is no
+ * provision for detecting further-removed ancestor processes, but if the
+ * init script is written carefully then all but the immediate parent
+ * shell will be root-owned processes and so the kill test will fail with
+ * EPERM. Note that we cannot get a false negative this way, because an
+ * existing postmaster would surely never launch a competing postmaster or
+ * pg_ctl process directly.
*/
my_pid = getpid();
@@ -1225,7 +1309,11 @@ CreateLockFile(const char *filename, bool amPostmaster,
*/
fd = open(filename, O_RDWR | O_CREAT | O_EXCL, pg_file_create_mode);
if (fd >= 0)
- break; /* Success; exit the retry loop */
+ {
+ /* Success; lock and exit the retry loop */
+ flock_fd = OFDLockFile(fd, filename);
+ break;
+ }
/*
* Couldn't create the pid file. Probably it already exists.
@@ -1239,8 +1327,12 @@ CreateLockFile(const char *filename, bool amPostmaster,
/*
* Read the file to get the old owner's PID. Note race condition
* here: file might have been deleted since we tried to create it.
+ *
+ * We're going to use the same fd for flock, and want to create a
+ * write lock for the latter one. Since both fd and the lock have to
+ * be of the same type, open the file for read and write.
*/
- fd = open(filename, O_RDONLY, pg_file_create_mode);
+ fd = open(filename, O_RDWR, pg_file_create_mode);
if (fd < 0)
{
if (errno == ENOENT)
@@ -1250,6 +1342,10 @@ CreateLockFile(const char *filename, bool amPostmaster,
errmsg("could not open lock file \"%s\": %m",
filename)));
}
+
+ /* Try to lock the file. We stop here, if it's already locked. */
+ flock_fd = OFDLockFile(fd, filename);
+
pgstat_report_wait_start(WAIT_EVENT_LOCK_FILE_CREATE_READ);
if ((len = read(fd, buffer, sizeof(buffer) - 1)) < 0)
ereport(FATAL,
@@ -1449,7 +1545,11 @@ CreateLockFile(const char *filename, bool amPostmaster,
* Use lcons so that the lock files are unlinked in reverse order of
* creation; this is critical!
*/
- lock_files = lcons(pstrdup(filename), lock_files);
+ lock_file = palloc0_object(PGLockFile);
+ lock_file->filename = pstrdup(filename);
+ lock_file->fd = flock_fd;
+
+ lock_files = lcons(lock_file, lock_files);
}
/*
@@ -1496,14 +1596,14 @@ TouchSocketLockFiles(void)
foreach(l, lock_files)
{
- char *socketLockFile = (char *) lfirst(l);
+ PGLockFile *lock_file = (PGLockFile *) lfirst(l);
/* No need to touch the data directory lock file, we trust */
- if (strcmp(socketLockFile, DIRECTORY_LOCK_FILE) == 0)
+ if (strcmp(lock_file->filename, DIRECTORY_LOCK_FILE) == 0)
continue;
/* we just ignore any error here */
- (void) utime(socketLockFile, NULL);
+ (void) utime(lock_file->filename, NULL);
}
}
diff --git a/src/include/pg_config.h.in b/src/include/pg_config.h.in
index 661c4a9b168..fae00af3284 100644
--- a/src/include/pg_config.h.in
+++ b/src/include/pg_config.h.in
@@ -85,6 +85,10 @@
don't. */
#undef HAVE_DECL_F_FULLFSYNC
+/* Define to 1 if you have the declaration of `F_OFD_SETLK', and to 0 if you
+ don't. */
+#undef HAVE_DECL_F_OFD_SETLK
+
/* Define to 1 if you have the declaration of `memset_s', and to 0 if you
don't. */
#undef HAVE_DECL_MEMSET_S
diff --git a/src/tools/pgindent/typedefs.list b/src/tools/pgindent/typedefs.list
index 59cf40b5bb0..fdd2ff2cc32 100644
--- a/src/tools/pgindent/typedefs.list
+++ b/src/tools/pgindent/typedefs.list
@@ -1958,6 +1958,7 @@ PGIOAlignedBlock
PGLZ_HistEntry
PGLZ_Strategy
PGLoadBalanceType
+PGLockFile
PGMessageField
PGModuleMagicFunction
PGNoticeHooks
base-commit: 45817676632b46fc9d2b97c8d432348e0ee4f1f4
--
2.55.0
^ permalink raw reply [nested|flat] 10+ messages in thread
end of thread, other threads:[~2026-08-28 07:28 UTC | newest]
Thread overview: 10+ messages (download: mbox mbox.gz follow: Atom feed)
-- links below jump to the message on this page --
2025-12-19 14:27 File locks for data directory lockfile in the context of Linux namespaces Dmitry Dolgov <9erthalion6@gmail.com>
2026-01-17 15:26 ` Dmitry Dolgov <9erthalion6@gmail.com>
2026-06-05 12:37 ` Ilmar Yunusov <tanswis42@gmail.com>
2026-06-19 15:11 ` Dmitry Dolgov <9erthalion6@gmail.com>
2026-06-23 14:29 ` Dmitry Dolgov <9erthalion6@gmail.com>
2026-07-06 07:30 ` Ilmar Yunusov <tanswis42@gmail.com>
2026-07-06 12:10 ` Dmitry Dolgov <9erthalion6@gmail.com>
2026-07-06 22:07 ` Zsolt Parragi <zsolt.parragi@percona.com>
2026-07-10 10:20 ` Dmitry Dolgov <9erthalion6@gmail.com>
2026-08-28 07:28 ` Dmitry Dolgov <9erthalion6@gmail.com>
This inbox is served by agora; see mirroring instructions
for how to clone and mirror all data and code used for this inbox