agora inbox for pgsql-hackers@postgresql.org  
help / color / mirror / Atom feed
From: Nathan Bossart <nathandbossart@gmail.com>
To: Nazir Bilal Yavuz <byavuz81@gmail.com>
Cc: Andrew Dunstan <andrew@dunslane.net>
Cc: KAZAR Ayoub <ma_kazar@esi.dz>
Cc: Shinya Kato <shinya11.kato@gmail.com>
Cc: pgsql-hackers@postgresql.org
Subject: Re: Speed up COPY FROM text/CSV parsing using SIMD
Date: Wed, 22 Oct 2025 14:24:59 -0500
Message-ID: <aPkvi5P7kpA8oQKc@nathan> (raw)
In-Reply-To: <CAN55FZ0AYP4ZEczBJ5ur-=9QuEhMysH9Yfrq5srr0ZakK1M0FA@mail.gmail.com>
References: <CAN55FZ1J+6eM=F5GreWEBMJcNV_gifYyYY1b6xpYzun=nWPhMQ@mail.gmail.com>
	<CAN55FZ109W90Ux_EBEqkkU2TyNqBNhdhN_1XPRGo3iiZ2L9b=A@mail.gmail.com>
	<8615c983-1662-43b4-b0c9-49d194ac33aa@dunslane.net>
	<CAN55FZ1KF7XNpm2XyG=M-sFUODai=6Z8a11xE3s4YRBeBKY3tA@mail.gmail.com>
	<673d92f7-2489-475f-a208-9414ea35d4d8@dunslane.net>
	<aPZrg6lxb5bgy_px@nathan>
	<8e045899-2023-48b1-bd91-f8cdffeb511d@dunslane.net>
	<CAN55FZ2GonAeSJHn-c2nJgUO-v6sDMOQzn97evVdZbcHeu3ihw@mail.gmail.com>
	<aPfTiX0HwV42R6Od@nathan>
	<CAN55FZ0AYP4ZEczBJ5ur-=9QuEhMysH9Yfrq5srr0ZakK1M0FA@mail.gmail.com>

On Wed, Oct 22, 2025 at 03:33:37PM +0300, Nazir Bilal Yavuz wrote:
> On Tue, 21 Oct 2025 at 21:40, Nathan Bossart <nathandbossart@gmail.com> wrote:
>> I wonder if we could mitigate the regression further by spacing out the
>> checks a bit more.  It could be worth comparing a variety of values to
>> identify what works best with the test data.
> 
> Do you mean that instead of doubling the SIMD sleep, we should
> multiply it by 3 (or another factor)? Or are you referring to
> increasing the maximum sleep from 1024? Or possibly both?

I'm not sure of the precise details, but the main thrust of my suggestion
is to assume that whatever sampling you do to determine whether to use SIMD
is good for a larger chunk of data.  That is, if you are sampling 1K lines
and then using the result to choose whether to use SIMD for the next 100K
lines, we could instead bump the latter number to 1M lines (or something).
That way we minimize the regression for relatively uniform data sets while
retaining some ability to adapt in case things change halfway through a
large table.

-- 
nathan





view thread (178+ messages)  latest in thread

Message-ID: <aPkvi5P7kpA8oQKc@nathan>
Permalink:  ../aPkvi5P7kpA8oQKc@nathan/
Also on:    postgresql.org/message-id/aPkvi5P7kpA8oQKc@nathan

reply

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Reply to all the recipients using the --to and --cc options:
  reply via email

  To: pgsql-hackers@postgresql.org
  Cc: nathandbossart@gmail.com, byavuz81@gmail.com, andrew@dunslane.net, ma_kazar@esi.dz, shinya11.kato@gmail.com
  Subject: Re: Speed up COPY FROM text/CSV parsing using SIMD
  In-Reply-To: <aPkvi5P7kpA8oQKc@nathan>

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

This inbox is served by agora; see mirroring instructions
for how to clone and mirror all data and code used for this inbox