agora inbox for [email protected]
help / color / mirror / Atom feedRe: measuring lwlock-related latency spikes
267+ messages / 3 participants
[nested] [flat]
* Re: measuring lwlock-related latency spikes
@ 2012-04-03 12:28 Kevin Grittner <[email protected]>
0 siblings, 1 reply; 267+ messages in thread
From: Kevin Grittner @ 2012-04-03 12:28 UTC (permalink / raw)
To: [email protected]; +Cc: [email protected]; [email protected]; pgsql-hackers
> Robert Haas wrote:
> Kevin Grittner wrote:
>
>> I can't help thinking that the "background hinter" I had ideas
>> about writing would prevent many of the reads of old CLOG pages,
>> taking a lot of pressure off of this area. It just occurred to me
>> that the difference between that idea and having an autovacuum
>> thread which just did first-pass work on dirty heap pages is slim
>> to none.
>
> Yeah. Marking things all-visible in the background seems possibly
> attractive, too. I think the trick is to figuring out the control
> mechanism. In this case, the workload fits within shared_buffers,
> so it's not helpful to think about using buffer eviction as the
> trigger for doing these operations, though that might have some
> legs in general. And a simple revolving scan over shared_buffers
> doesn't really figure to work out well either, I suspect, because
> it's too undirected. I think what you'd really like to have is a
> list of buffers that were modified by transactions which have
> recently committed or rolled back.
Yeah, that's what I was thinking. Since we only care about dirty
unhinted tuples, we need some fairly efficient way to track those to
make this pay.
> but that seems more like a nasty benchmarking kludge that something
> that's likely to solve real-world problems.
I'm not so sure. Unfortunagely, it may be hard to know without
writing at least a crude form of this to test, but there are several
workloads where hint bit rewrites and/or CLOG contention caused by
the slow tapering of usage of old pages contribute to problems.
>> I know how much time good benchmarking can take, so I hesitate to
>> suggest another permutation, but it might be interesting to see
>> what it does to the throughput if autovacuum is configured to what
>> would otherwise be considered insanely aggressive values (just for
>> vacuum, not analyze). To give this a fair shot, the whole database
>> would need to be vacuumed between initial load and the start of
>> the benchmark.
>
> If you would like to provide a chunk of settings that I can splat
> into postgresql.conf, I'm happy to run 'em through a test cycle and
> see what pops out.
Might as well jump in with both feet:
autovacuum_naptime = 1s
autovacuum_vacuum_threshold = 1
autovacuum_vacuum_scale_factor = 0.0
If that smooths the latency peaks and doesn't hurt performance too
much, it's decent evidence that the more refined technique could be a
win.
-Kevin
^ permalink raw reply [nested|flat] 267+ messages in thread
* Re: measuring lwlock-related latency spikes
@ 2012-04-06 04:30 Robert Haas <[email protected]>
parent: Kevin Grittner <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Robert Haas @ 2012-04-06 04:30 UTC (permalink / raw)
To: Kevin Grittner <[email protected]>; +Cc: [email protected]; [email protected]; pgsql-hackers
On Tue, Apr 3, 2012 at 8:28 AM, Kevin Grittner
<[email protected]> wrote:
> Might as well jump in with both feet:
>
> autovacuum_naptime = 1s
> autovacuum_vacuum_threshold = 1
> autovacuum_vacuum_scale_factor = 0.0
>
> If that smooths the latency peaks and doesn't hurt performance too
> much, it's decent evidence that the more refined technique could be a
> win.
It seems this isn't good for either throughput or latency. Here are
latency percentiles for a recent run against master with my usual
settings:
90 1668
91 1747
92 1845
93 1953
94 2064
95 2176
96 2300
97 2461
98 2739
99 3542
100 12955473
And here's how it came out with these settings:
90 1818
91 1904
92 1998
93 2096
94 2200
95 2316
96 2459
97 2660
98 3032
99 3868
100 10842354
tps came out tps = 13658.330709 (including connections establishing),
vs 14546.644712 on the other run.
I have a (possibly incorrect) feeling that even with these
ridiculously aggressive settings, nearly all of the cleanup work is
getting done by HOT prunes rather than by vacuum, so we're still not
testing what we really want to be testing, but we're doing a lot of
extra work along the way.
--
Robert Haas
EnterpriseDB: http://www.enterprisedb.com
The Enterprise PostgreSQL Company
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
* [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation.
@ 2020-06-12 02:38 Andres Freund <[email protected]>
0 siblings, 0 replies; 267+ messages in thread
From: Andres Freund @ 2020-06-12 02:38 UTC (permalink / raw)
---
src/include/portability/instr_time.h | 68 ++++++++++++++++++++++++----
1 file changed, 60 insertions(+), 8 deletions(-)
diff --git a/src/include/portability/instr_time.h b/src/include/portability/instr_time.h
index fc058d548a8..8b2f9a2e707 100644
--- a/src/include/portability/instr_time.h
+++ b/src/include/portability/instr_time.h
@@ -83,7 +83,9 @@
#define PG_INSTR_CLOCK CLOCK_REALTIME
#endif
+/* time in baseline cpu cycles */
typedef int64 instr_time;
+
#define NS_PER_S INT64CONST(1000000000)
#define US_PER_S INT64CONST(1000000)
#define MS_PER_S INT64CONST(1000)
@@ -95,17 +97,67 @@ typedef int64 instr_time;
#define INSTR_TIME_SET_ZERO(t) ((t) = 0)
-static inline instr_time pg_clock_gettime_ns(void)
+#include <x86intrin.h>
+#include <cpuid.h>
+
+/*
+ * Return what the number of cycles needs to be multiplied with to end up with
+ * seconds.
+ *
+ * FIXME: The cold portion should probably be out-of-line. And it'd be better
+ * to not recompute this in every file that uses this. Best would probably be
+ * to require explicit initialization of cycles_to_sec, because having a
+ * branch really is unnecessary.
+ *
+ * FIXME: We should probably not unnecessarily use floating point math
+ * here. And it's likely that the numbers are small enough that we are running
+ * into floating point inaccuracies already. Probably worthwhile to be a good
+ * bit smarter.
+ *
+ * FIXME: This would need to be conditional, with a fallback to something not
+ * rdtsc based.
+ */
+static inline double __attribute__((const))
+get_cycles_to_sec(void)
{
- struct timespec tmp;
+ static double cycles_to_sec = 0;
- clock_gettime(PG_INSTR_CLOCK, &tmp);
+ /*
+ * Compute baseline cpu peformance, determines speed at which rdtsc advances
+ */
+ if (unlikely(cycles_to_sec == 0))
+ {
+ uint32 cpuinfo[4] = {0};
- return tmp.tv_sec * NS_PER_S + tmp.tv_nsec;
+ __get_cpuid(0x16, cpuinfo, cpuinfo + 1, cpuinfo + 2, cpuinfo + 3);
+ cycles_to_sec = 1 / ((double) cpuinfo[0] * 1000 * 1000);
+ }
+
+ return cycles_to_sec;
+}
+
+static inline instr_time pg_clock_gettime_ref_cycles(void)
+{
+ /*
+ * The rdtscp waits for all in-flight instructions to finish (but allows
+ * later instructions to start concurrently). That's good for some timing
+ * situations (when the time is supposed to cover all the work), but
+ * terrible for others (when sub-parts of work are measured, because then
+ * the pipeline stall due to the wait change the overall timing).
+ */
+#if 0
+ unsigned int aux;
+ int64 tsc = __rdtscp(&aux);
+
+ return tsc;
+#else
+
+ return __rdtsc();
+#endif
}
#define INSTR_TIME_SET_CURRENT(t) \
- (t) = pg_clock_gettime_ns()
+ (t) = pg_clock_gettime_ref_cycles()
#define INSTR_TIME_ADD(x,y) \
do { \
@@ -123,13 +175,13 @@ static inline instr_time pg_clock_gettime_ns(void)
} while (0)
#define INSTR_TIME_GET_DOUBLE(t) \
- ((double) (t) / NS_PER_S)
+ ((double) (t) * get_cycles_to_sec())
#define INSTR_TIME_GET_MILLISEC(t) \
- ((double) (t) / NS_PER_MS)
+ ((double) (t) * (get_cycles_to_sec() * MS_PER_S))
#define INSTR_TIME_GET_MICROSEC(t) \
- ((double) (t) / NS_PER_US)
+ ((double) (t) * (get_cycles_to_sec() * US_PER_S))
#else /* !HAVE_CLOCK_GETTIME */
--
2.25.0.114.g5b0ca878e0
--wn4ncs637ccpvbhb--
^ permalink raw reply [nested|flat] 267+ messages in thread
end of thread, other threads:[~2020-06-12 02:38 UTC | newest]
Thread overview: 267+ messages (download: mbox mbox.gz follow: Atom feed)
-- links below jump to the message on this page --
2012-04-03 12:28 Re: measuring lwlock-related latency spikes Kevin Grittner <[email protected]>
2012-04-06 04:30 ` Robert Haas <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
2020-06-12 02:38 [PATCH v1 2/2] WIP: Use cpu reference cycles, via rdtsc, to measure time for instrumentation. Andres Freund <[email protected]>
This inbox is served by agora; see mirroring instructions
for how to clone and mirror all data and code used for this inbox