Received: from malur.postgresql.org ([217.196.149.56]) by arkaria.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.94.2) (envelope-from ) id 1sjpDG-004QkD-N2 for pgsql-hackers@arkaria.postgresql.org; Fri, 30 Aug 2024 00:10:02 +0000 Received: from localhost ([127.0.0.1] helo=malur.postgresql.org) by malur.postgresql.org with esmtp (Exim 4.94.2) (envelope-from ) id 1sjpDE-00AEEB-QY for pgsql-hackers@arkaria.postgresql.org; Fri, 30 Aug 2024 00:10:01 +0000 Received: from magus.postgresql.org ([2a02:c0:301:0:ffff::29]) by malur.postgresql.org with esmtps (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.94.2) (envelope-from ) id 1sjpDD-00AEC6-UV for pgsql-hackers@lists.postgresql.org; Fri, 30 Aug 2024 00:10:01 +0000 Received: from m16.mail.163.com ([117.135.210.2]) by magus.postgresql.org with esmtp (Exim 4.94.2) (envelope-from ) id 1sjpD8-002ABe-6E for pgsql-hackers@postgresql.org; Fri, 30 Aug 2024 00:09:59 +0000 DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=163.com; s=s110527; h=From:Subject:Date:Message-ID:MIME-Version: Content-Type; bh=hiR7Re0qYG5Zi9OMCx5yUECVrsaT5fYIs9sZ+069yK4=; b=QQhUDA83On1/q4xsalhlUzVxcn0vz4LubCQ9LkXZSiqKGF7gF8bBNSqxF1hReA QtzhqfjZfRjaNmZHXIlqMIhD+cQ+98iWM2k7yusVFLKZn1PjZVPWxkSSrOFQghJO 2yvdXErNFahu9unv1FlSk64bH7kMoTILUq3LbtPiBHOXI= Received: from lovely-coding (unknown [101.227.46.166]) by gzga-smtp-mta-g3-0 (Coremail) with SMTP id _____wDH77XHDdFm9vNvBA--.39740S3; Fri, 30 Aug 2024 08:09:44 +0800 (CST) From: Andy Fan To: David Rowley Cc: pgsql-hackers ,Tom Lane Subject: Re: Make printtup a bit faster In-Reply-To: (David Rowley's message of "Thu, 29 Aug 2024 23:51:48 +1200") References: <87wmjzfz0h.fsf@163.com> Date: Fri, 30 Aug 2024 08:09:43 +0800 Message-ID: <87bk1aj2go.fsf@163.com> MIME-Version: 1.0 Content-Type: text/plain X-CM-TRANSID: _____wDH77XHDdFm9vNvBA--.39740S3 X-Coremail-Antispam: 1Uf129KBjvJXoW7WFWftF4UCw1rGw4ktr1DJrb_yoW8Kr15pa 9Iy34ak3yDCas2ywn7AF4fX34fCr4UXrW3CFn5W3y0v398WFn7tF4akr4jkw17Wrn5Gw1Y qa1jv3Z8Kan3Zr7anT9S1TB71UUUUU7qnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDUYxBIdaVFxhVjvjDU0xZFpf9x0zRrR6xUUUUU= X-Originating-IP: [101.227.46.166] X-CM-SenderInfo: x2klx3xlid0iqsrtqiywtou0bp/1tbiNg1LU2XAnPrnWQAAse List-Id: List-Help: List-Subscribe: List-Post: List-Owner: List-Archive: Archived-At: Precedence: bulk David Rowley writes: Hello David, >> My high level proposal is define a type specific print function like: >> >> oidprint(Datum datum, StringInfo buf) >> textprint(Datum datum, StringInfo buf) > > I think what we should do instead is make the output functions take a > StringInfo and just pass it the StringInfo where we'd like the bytes > written. > > That of course would require rewriting all the output functions for > all the built-in types, so not a small task. Extensions make that job > harder. I don't think it would be good to force extensions to rewrite > their output functions, so perhaps some wrapper function could help us > align the APIs for extensions that have not been converted yet. I have the similar concern as Tom that this method looks too aggressive. That's why I said: "If a type's print function is not defined, we can still using the out function." AND "Hard coding the relationship between [common] used type and {type}print function OID looks not cool, Adding a new attribute in pg_type looks too aggressive however. Anyway this is the next topic to talk about." What would be the extra benefit we redesign all the out functions? > There's a similar problem with input functions not having knowledge of > the input length. You only have to look at textin() to see how useful > that could be. Fixing that would probably make COPY FROM horrendously > faster. Team that up with SIMD for the delimiter char search and COPY > go a bit better still. Neil Conway did propose the SIMD part in [1], > but it's just not nearly as good as it could be when having to still > perform the strlen() calls. OK, I think I can understand the needs to make in-function knows the input length and good to know the SIMD part for delimiter char search. strlen looks like a delimiter char search ('\0') as well. Not sure if "strlen" has been implemented with SIMD part, but if not, why? > I had planned to work on this for PG18, but I'd be happy for some > assistance if you're willing. I see you did many amazing work with cache-line-frindly data struct design, branch predition optimization and SIMD optimization. I'd like to try one myself. I'm not sure if I can meet the target, what if we handle the out/in function separately (can be by different people)? -- Best Regards Andy Fan