Received: from malur.postgresql.org ([217.196.149.56]) by arkaria.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.96) (envelope-from ) id 1w138l-002Vq9-0W for pgsql-hackers@arkaria.postgresql.org; Fri, 13 Mar 2026 14:05:23 +0000 Received: from localhost ([127.0.0.1] helo=malur.postgresql.org) by malur.postgresql.org with esmtp (Exim 4.96) (envelope-from ) id 1w138i-004Vml-28 for pgsql-hackers@arkaria.postgresql.org; Fri, 13 Mar 2026 14:05:21 +0000 Received: from makus.postgresql.org ([2001:4800:3e1:1::229]) by malur.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.96) (envelope-from ) id 1w138i-004Vmd-0o for pgsql-hackers@lists.postgresql.org; Fri, 13 Mar 2026 14:05:20 +0000 Received: from mail-oi1-x22d.google.com ([2607:f8b0:4864:20::22d]) by makus.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_128_GCM_SHA256 (Exim 4.98.2) (envelope-from ) id 1w138g-00000001xjm-0DkX for pgsql-hackers@postgresql.org; Fri, 13 Mar 2026 14:05:19 +0000 Received: by mail-oi1-x22d.google.com with SMTP id 5614622812f47-46708149af2so1246754b6e.0 for ; Fri, 13 Mar 2026 07:05:17 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1773410716; x=1774015516; darn=postgresql.org; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc:subject:date:message-id:reply-to; bh=fito5iEeXDxvGrgko9tJV4XGTmTPI+48Qt3Np4L9Kjo=; b=ZCLIJJLC/Cal+q2iuUJgzzgwJg1FLcn6/vTUmsb5oLSX3xd4vFnvlj45zT/fsYtaDj iW2Y3hPuNNe4urrtmie1ui4Q8kCJi1ZTYT07VLNTNjgKs+D6MfFhizgdASJuo5VuygU7 wutYJ0LsUFvIQqR7OV3lQ2lFrCFMVCI3fXtohutg3SQuC9wyI3s0+i4wuQ59s7CDzzEs md7c5LGelPOUOdgEBsDrrZVWemyrqeoVM/dkuem4JT8iYqyGO0y6nVwjyzmUd7/BWcN5 SJyP6FBfgpoh7UuGMlaDF84z0RYh9YK62rQmhhizdfoQpqx8NxRMiJ+D+7I35vvObl5n xxRQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1773410716; x=1774015516; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-gg:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=fito5iEeXDxvGrgko9tJV4XGTmTPI+48Qt3Np4L9Kjo=; b=MDX0iCBwx2rdCztqHEoDEHKJTuU5rpgyjn0TciI6bMgGEtYtCTFl/2j62uh9JGM0PB tXmfVqF/6DArZjrNe2Riop3p0kBaA1tiM4NmUsXLdpdCOxJiqYk1rmwubVe6/9AXaItR 4GIFzhM+HNEXx8NGNvJRSHvkdl1xkLcQ+SqG/o6F/ifkiJKwHcyb7ZVBU+Mqpilf4q+S YRrLp95CQ30JC1CljhoYauwwj0s2cRxjriZ1s74Xbs9TBVC3u8KOVaYdj7+h02xiT1me S33FRjAFuDK+eQ80esdtgfX0UPvpKadXfCgyd8MarTkfSdfLMOza/wdT1rEgolpvgGbN QeUg== X-Forwarded-Encrypted: i=1; AJvYcCVZwzzdvnQHlOo+9x6pPpIb4Tbzq6OZtz7HPjDOGaeImSPUBNyj+0c+tWKTIL2/8YRndYyIaeV19S1m0EuF@postgresql.org X-Gm-Message-State: AOJu0Yx4BqV2gc8kV7MfK85EruioEzSv4VStIYYT8nZ9KZ0a0x4LmobJ mAMs8umk7Q0USgjj5dF+DdpRwXKahXpEDzbAbYycO3gt3FAeGVD9W50s X-Gm-Gg: ATEYQzxjGp8egcEIi3NTA5anZax58pm/u5fSbU3mOPtM2TvccnJ8LiKvy3+VkTC2SLr qknUOfmDi+X0ePMt4MY+5LUFp0A434llIbJfD7janEW1jNYQXRGNbIy7QcVj1aM0P4cZBouQ4x3 YPhxAswCjwm87mYmpH1h8aBEbon6YEYRx05q4Gy1ujIN9GcMVydGh/TzcqaUyb/NdEsJOw0K5O/ uZV4ojtYMDXxvyTj1XjQD/v7KxSMmbiWCUEtE9+eZE3BxWXp9+uV+1jVMOPeiQImlt2jayjhBzO wDQRguQuFn+E8Ti95abjX5LgkhbdHG5S8HKKc/dHiZDFUpHWGaw0fdFVNU2kIo70KMgEx0MzLB8 D94fbdRVIBNrAuvQOUK8BeHUE+Mbxq9SxPglgW34h+2pMCQdzDDWQ0Ww6c0s8eo3KNvemx89uXt Jtv0Lm18BdESlBuUakt8LNv1ooqn1SuJGYgejZ7WXLYg9gZjCNcE1tXV2nw0vAYv7oGiiyJyPUE aYU3Lb6vjEVcJnv1gMezQ== X-Received: by 2002:a05:6808:2f17:b0:467:1212:46fe with SMTP id 5614622812f47-467572d39f7mr1851611b6e.38.1773410715906; Fri, 13 Mar 2026 07:05:15 -0700 (PDT) Received: from nathan (162-195-168-172.lightspeed.stlsmo.sbcglobal.net. [162.195.168.172]) by smtp.gmail.com with ESMTPSA id 5614622812f47-467342ef853sm4840234b6e.15.2026.03.13.07.05.14 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 13 Mar 2026 07:05:14 -0700 (PDT) Date: Fri, 13 Mar 2026 09:05:13 -0500 From: Nathan Bossart To: Nazir Bilal Yavuz Cc: Manni Wood , KAZAR Ayoub , Neil Conway , Andrew Dunstan , Shinya Kato , PostgreSQL-development Subject: Re: Speed up COPY FROM text/CSV parsing using SIMD Message-ID: References: MIME-Version: 1.0 Content-Type: multipart/mixed; boundary="3SATdDqw4zXzoVGn" Content-Disposition: inline In-Reply-To: List-Id: List-Help: List-Subscribe: List-Post: List-Owner: List-Archive: Archived-At: Precedence: bulk --3SATdDqw4zXzoVGn Content-Type: text/plain; charset=us-ascii Content-Disposition: inline On Fri, Mar 13, 2026 at 04:34:49PM +0300, Nazir Bilal Yavuz wrote: > On Fri, 13 Mar 2026 at 14:57, Nazir Bilal Yavuz wrote: >> Unfortunately, v15 causes a regression for a 'csv & wide & 1/3' case >> on my end. v14 was taking 8000ms but v15 took ~9100ms. If we add the >> tmp_hit_eof variable then the regression disappears. Also, if I use a >> struct like below, regression disappears again. > >> When I removed the tmp_hit_eof variable on v14, I didn't encounter any >> regression. I really don't understand why this is happening on my end. >> Manni didn't encounter any regression on the benchmark [1]. > > Problem might be related to gcc. I am using Debian Trixie and my > current gcc version is 'gcc version 14.2.0 (Debian 14.2.0-19)'. If I > compile Postgres with 'Debian clang version 19.1.7 (3+b1)', then there > is no regression, which makes more sense IMO. Let's just re-add the temporary variable for hit_eof. The struct idea is clever, but it's just a little more complicated than I think is necessary here. I've also removed the goto in favor of just duplicating the "out" code, like you had before. I'd like to avoid sporadic #ifndef USE_NO_SIMD uses, and goto is out of fashion, anyway. -- nathan --3SATdDqw4zXzoVGn Content-Type: text/plain; charset=us-ascii Content-Disposition: attachment; filename=v17-0001-Optimize-COPY-FROM-FORMAT-text-csv-using-SIMD.patch From cdb600ff363a583e53594e03f459e3ab3eea579e Mon Sep 17 00:00:00 2001 From: Nathan Bossart Date: Thu, 12 Mar 2026 12:32:23 -0500 Subject: [PATCH v17 1/1] Optimize COPY FROM (FORMAT {text,csv}) using SIMD. Presently, such commands scan the input buffer one byte at a time looking for special characters. This commit adds a new path that uses SIMD instructions to skip over chunks of data without any special characters. This can be much faster. To avoid regressions, SIMD processing is disabled for the remainder of the COPY FROM command as soon as we encounter a short line or a special character (except for end-of-line characters, else we'd always disable it after the first line). This is perhaps too conservative, but it could probably be made more lenient in the future via fine-tuned heuristics. Author: Nazir Bilal Yavuz Co-authored-by: Shinya Kato Reviewed-by: Ayoub Kazar Reviewed-by: Andrew Dunstan Reviewed-by: Neil Conway Tested-by: Manni Wood Tested-by: Mark Wong Discussion: https://postgr.es/m/CAOzEurSW8cNr6TPKsjrstnPfhf4QyQqB4tnPXGGe8N4e_v7Jig%40mail.gmail.com --- src/backend/commands/copyfrom.c | 1 + src/backend/commands/copyfromparse.c | 185 ++++++++++++++++++++++- src/include/commands/copyfrom_internal.h | 1 + 3 files changed, 184 insertions(+), 3 deletions(-) diff --git a/src/backend/commands/copyfrom.c b/src/backend/commands/copyfrom.c index 0ece40557c8..95f6cb416a9 100644 --- a/src/backend/commands/copyfrom.c +++ b/src/backend/commands/copyfrom.c @@ -1746,6 +1746,7 @@ BeginCopyFrom(ParseState *pstate, cstate->cur_attname = NULL; cstate->cur_attval = NULL; cstate->relname_only = false; + cstate->simd_enabled = true; /* * Allocate buffers for the input pipeline. diff --git a/src/backend/commands/copyfromparse.c b/src/backend/commands/copyfromparse.c index 84c8809a889..00ee4154b8b 100644 --- a/src/backend/commands/copyfromparse.c +++ b/src/backend/commands/copyfromparse.c @@ -72,6 +72,7 @@ #include "miscadmin.h" #include "pgstat.h" #include "port/pg_bswap.h" +#include "port/simd.h" #include "utils/builtins.h" #include "utils/rel.h" #include "utils/wait_event.h" @@ -1311,6 +1312,152 @@ CopyReadLine(CopyFromState cstate, bool is_csv) return result; } +#ifndef USE_NO_SIMD +/* + * Helper function for CopyReadLineText() that uses SIMD instructions to scan + * the input buffer for special characters. This can be much faster. + * + * Note that we disable SIMD for the remainder of the COPY FROM command upon + * encountering a special character (except for end-of-line characters) or a + * short line. This is perhaps too conservative, but it should help avoid + * regressions. It could probably be made more lenient in the future via + * fine-tuned heuristics. + */ +static bool +CopyReadLineTextSIMDHelper(CopyFromState cstate, bool is_csv, + bool *hit_eof_p, int *input_buf_ptr_p) +{ + char *copy_input_buf; + int input_buf_ptr; + int copy_buf_len; + bool unique_esc_char; /* for csv, do quote/esc chars differ? */ + bool first = true; + bool result = false; + const Vector8 nl_vec = vector8_broadcast('\n'); + const Vector8 cr_vec = vector8_broadcast('\r'); + Vector8 bs_or_quote_vec; /* '\' for text, quote for csv */ + Vector8 esc_vec; /* only for csv */ + + if (is_csv) + { + char quote = cstate->opts.quote[0]; + char esc = cstate->opts.escape[0]; + + bs_or_quote_vec = vector8_broadcast(quote); + esc_vec = vector8_broadcast(esc); + unique_esc_char = (quote != esc); + } + else + { + bs_or_quote_vec = vector8_broadcast('\\'); + unique_esc_char = false; + } + + /* + * For a little extra speed within the loop, we copy some state members + * into local variables. Note that we need to use a separate local + * variable for input_buf_ptr so that the REFILL_LINEBUF macro works. We + * copy its value into the input_buf_ptr_p argument before returning. + */ + copy_input_buf = cstate->input_buf; + input_buf_ptr = cstate->input_buf_index; + copy_buf_len = cstate->input_buf_len; + + /* + * See the corresponding loop in CopyReadLineText() for more information + * about the purpose of this loop. This one does the same thing using + * SIMD instructions, although we are quick to bail out to the scalar path + * if we encounter a special character. + */ + for (;;) + { + Vector8 chunk; + Vector8 match; + + /* Load more data if needed. */ + if (copy_buf_len - input_buf_ptr < sizeof(Vector8)) + { + REFILL_LINEBUF; + + CopyLoadInputBuf(cstate); + /* update our local variables */ + *hit_eof_p = cstate->input_reached_eof; + input_buf_ptr = cstate->input_buf_index; + copy_buf_len = cstate->input_buf_len; + + /* + * If we are completely out of data, break out of the loop, + * reporting EOF. + */ + if (INPUT_BUF_BYTES(cstate) <= 0) + { + result = true; + break; + } + } + + /* + * If we still don't have enough data for the SIMD path, fall back to + * the scalar code. Note that this doesn't necessarily mean we + * encountered a short line, so we leave cstate->simd_enabled set to + * true. + */ + if (copy_buf_len - input_buf_ptr < sizeof(Vector8)) + break; + + /* + * If we made it here, we have at least enough data to fit in a + * Vector8, so we can use SIMD instructions to scan for special + * characters. + */ + vector8_load(&chunk, (const uint8 *) ©_input_buf[input_buf_ptr]); + + /* + * Check for \n, \r, \\ (for text), quotes (for csv), and escapes (for + * csv, if different from quotes). + */ + match = vector8_eq(chunk, nl_vec); + match = vector8_or(match, vector8_eq(chunk, cr_vec)); + match = vector8_or(match, vector8_eq(chunk, bs_or_quote_vec)); + if (unique_esc_char) + match = vector8_or(match, vector8_eq(chunk, esc_vec)); + + /* + * If we found a special character, advance to it and hand off to the + * scalar path. Except for end-of-line characters, we also disable + * SIMD processing for the remainder of the COPY FROM command. + */ + if (vector8_is_highbit_set(match)) + { + uint32 mask; + char c; + + mask = vector8_highbit_mask(match); + input_buf_ptr += pg_rightmost_one_pos32(mask); + + /* + * Don't disable SIMD if we found \n or \r, else we'd stop using + * SIMD instructions after the first line. As an exception, we do + * disable it if this is the first vector we processed, as that + * means the line is too short for SIMD. + */ + c = copy_input_buf[input_buf_ptr]; + if (first || (c != '\n' && c != '\r')) + cstate->simd_enabled = false; + + break; + } + + /* That chunk was clear of special characters, so we can skip it. */ + input_buf_ptr += sizeof(Vector8); + first = false; + } + + *input_buf_ptr_p = input_buf_ptr; + return result; +} +#endif /* ! USE_NO_SIMD */ + /* * CopyReadLineText - inner loop of CopyReadLine for text mode */ @@ -1361,11 +1508,43 @@ CopyReadLineText(CopyFromState cstate, bool is_csv) * input_buf_ptr have been determined to be part of the line, but not yet * transferred to line_buf. * - * For a little extra speed within the loop, we copy input_buf and - * input_buf_len into local variables. + * For a little extra speed within the loop, we copy some state + * information into local variables. input_buf_ptr could be changed in + * the SIMD path, so we must set that one before it. The others are set + * afterwards. */ - copy_input_buf = cstate->input_buf; input_buf_ptr = cstate->input_buf_index; + + /* + * We first try to use SIMD for the task described above, falling back to + * the scalar path (i.e., the loop below) if needed. + */ +#ifndef USE_NO_SIMD + if (cstate->simd_enabled) + { + /* + * Using temporary variables seems to encourage the compiler to keep + * them in a register, which is beneficial for performance. + */ + bool tmp_hit_eof = false; + int tmp_input_buf_ptr = 0; /* silence compiler warning */ + + result = CopyReadLineTextSIMDHelper(cstate, is_csv, &tmp_hit_eof, + &tmp_input_buf_ptr); + hit_eof = tmp_hit_eof; + input_buf_ptr = tmp_input_buf_ptr; + + if (result) + { + /* Transfer any still-uncopied data to line_buf. */ + REFILL_LINEBUF; + + return result; + } + } +#endif /* ! USE_NO_SIMD */ + + copy_input_buf = cstate->input_buf; copy_buf_len = cstate->input_buf_len; for (;;) diff --git a/src/include/commands/copyfrom_internal.h b/src/include/commands/copyfrom_internal.h index f892c343157..9d3e244ee55 100644 --- a/src/include/commands/copyfrom_internal.h +++ b/src/include/commands/copyfrom_internal.h @@ -108,6 +108,7 @@ typedef struct CopyFromStateData * att */ bool *defaults; /* if DEFAULT marker was found for * corresponding att */ + bool simd_enabled; /* use SIMD to scan for special chars? */ /* * True if the corresponding attribute's is a constrained domain. This -- 2.50.1 (Apple Git-155) --3SATdDqw4zXzoVGn--