Received: from malur.postgresql.org ([217.196.149.56]) by arkaria.postgresql.org with esmtps (TLS1.3:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.92) (envelope-from ) id 1oS42o-00076C-Jm for pgsql-hackers@arkaria.postgresql.org; Sat, 27 Aug 2022 22:12:46 +0000 Received: from localhost ([127.0.0.1] helo=malur.postgresql.org) by malur.postgresql.org with esmtp (Exim 4.92) (envelope-from ) id 1oS42m-0001ry-Bg for pgsql-hackers@arkaria.postgresql.org; Sat, 27 Aug 2022 22:12:44 +0000 Received: from magus.postgresql.org ([2a02:c0:301:0:ffff::29]) by malur.postgresql.org with esmtps (TLS1.3:ECDHE_RSA_AES_256_GCM_SHA384:256) (Exim 4.92) (envelope-from ) id 1oS42m-0001rQ-1F for pgsql-hackers@lists.postgresql.org; Sat, 27 Aug 2022 22:12:44 +0000 Received: from mail-pf1-x433.google.com ([2607:f8b0:4864:20::433]) by magus.postgresql.org with esmtps (TLS1.3:ECDHE_RSA_AES_128_GCM_SHA256:128) (Exim 4.92) (envelope-from ) id 1oS42h-0001ne-2o for pgsql-hackers@postgresql.org; Sat, 27 Aug 2022 22:12:43 +0000 Received: by mail-pf1-x433.google.com with SMTP id 142so4831327pfu.10 for ; Sat, 27 Aug 2022 15:12:38 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20210112; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:from:to:cc; bh=jMfEzEnsVgX2iwXVTekMGAYef8bEdOZXzm2h6Iq2vGk=; b=Lu9dukp0gn6CjcJbcNM1eH+szj5IJdNL+aaRNk8dr1JfgndUOxHqrgDjNMBYPbekNR rm4/K61Yr7A0B3fpYhpi6KwdSNUOGbhtw+/eMICoJp0cAen6D+Nh+m98yxNTrBgQeoH4 kpy7VPednsof2+hxZV/JhvRI3cbu/oXQrVRq7uCWA7upsPAZlfDLokGOnzcR1Lqh3ORj jJ++TEbPX7SW/8khomdPPgO7ppp6na02ZZXlL3f2aMNMQzqU10pjji3O5OEoHLiF95rl 0cLhwzeKD+EhqfKTkefGUK+M13EK3p8/F5awMOpJhs8b1IgSXuqYkc9GHvIdQRksX7ZY CXQw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20210112; h=in-reply-to:content-disposition:mime-version:references:message-id :subject:cc:to:from:date:x-gm-message-state:from:to:cc; bh=jMfEzEnsVgX2iwXVTekMGAYef8bEdOZXzm2h6Iq2vGk=; b=EvBmXBPxbyioOnJ71TxGpWso9XmcQCkPZn86zJlBMrU8FC/ZpDpVPmEXxI2WeGtZTq 7P2JWbt9dst4zPZJ7ac7VYTHumJO731FUFPtN8Y9eZzAcT7JGef0Exg97J59tpMEXFaM 2FcY7zEPaTTUjE1YHQ12G9LrKSV/vlLeeSHlG94C0tf4WGQJEZHzfmh/aIQvJFu+6PoI +mWNp6jjzr27RoJZPJk/Ya1ZGih5iDxs3hpPQguI92wzCRO1/ZpO/29AGlUcP7+TGoIW SkLQ8vn44tVpE+hEFR7onuMrtz/vTUWPKMcKdfFJyrZI+36Hx66Z+iSNbWmEC+BiOAxM zF6Q== X-Gm-Message-State: ACgBeo1LvPWbwH0WGjYgghk8OTiXWiJEX2KAVCjdnGaopmWjRU7GdtBX emyGG+R3EzOQnHT+VF/QbK0= X-Google-Smtp-Source: AA6agR4yxtnM+UTCG1WPjV5ajGqHTOEfeyMLS/jMnW3dq6PewPntwdbUVYUse29icONv+MJDV3DMsw== X-Received: by 2002:a63:5b10:0:b0:429:c287:7bfa with SMTP id p16-20020a635b10000000b00429c2877bfamr8058802pgb.347.1661638357028; Sat, 27 Aug 2022 15:12:37 -0700 (PDT) Received: from nathanxps13 ([50.47.162.83]) by smtp.gmail.com with ESMTPSA id u6-20020a170902e5c600b0017297a6b39dsm4162249plf.265.2022.08.27.15.12.35 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Sat, 27 Aug 2022 15:12:36 -0700 (PDT) Date: Sat, 27 Aug 2022 15:12:34 -0700 From: Nathan Bossart To: John Naylor Cc: Andres Freund , pgsql-hackers Subject: Re: use ARM intrinsics in pg_lfind32() where available Message-ID: <20220827221234.GA15951@nathanxps13> References: <20220822211547.GA1126462@nathanxps13> <20220824180111.GB1302810@nathanxps13> <20220825045729.GA1458024@nathanxps13> <20220826045115.GA1638993@nathanxps13> <20220826061347.GA1777731@nathanxps13> <20220826182403.GA1917683@nathanxps13> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: List-Id: List-Help: List-Subscribe: List-Post: List-Owner: List-Archive: Archived-At: Precedence: bulk Thanks for taking a look. On Sat, Aug 27, 2022 at 01:59:06PM +0700, John Naylor wrote: > I don't forsee any use of emulating vector registers with uint64 if > they only hold two ints. I wonder if it'd be better if all vector32 > functions were guarded with #ifndef NO_USE_SIMD. (I wonder if > declarations without definitions cause warnings...) Yeah. I was a bit worried about the readability of this file with so many #ifndefs, but after trying it out, I suppose it doesn't look _too_ bad. > + * NB: This function assumes that each lane in the given vector either has all > + * bits set or all bits zeroed, as it is mainly intended for use with > + * operations that produce such vectors (e.g., vector32_eq()). If this > + * assumption is not true, this function's behavior is undefined. > + */ > > Hmm? Yup. The problem is that AFAICT there's no equivalent to _mm_movemask_epi8() on aarch64, so you end up with something like vmaxvq_u8(vandq_u8(v, vector8_broadcast(0x80))) != 0 But for pg_lfind32(), we really just want to know if any lane is set, which only requires a call to vmaxvq_u32(). I haven't had a chance to look too closely, but my guess is that this ultimately results in an extra AND operation in the aarch64 path, so maybe it doesn't impact performance too much. The other option would be to open-code the intrinsic function calls into pg_lfind.h. I'm trying to avoid the latter, but maybe it's the right thing to do for now... What do you think? > -#elif defined(USE_SSE2) > +#elif defined(USE_SSE2) || defined(USE_NEON) > > I think we can just say #else. Yes. > -#if defined(USE_SSE2) > - __m128i sub; > +#ifndef USE_NO_SIMD > + Vector8 sub; > > +#elif defined(USE_NEON) > + > + /* use the same approach as the USE_SSE2 block above */ > + sub = vqsubq_u8(v, vector8_broadcast(c)); > + result = vector8_has_zero(sub); > > I think we should invent a helper that does saturating subtraction and > call that, inlining the sub var so we don't need to mess with it > further. Good idea, will do. -- Nathan Bossart Amazon Web Services: https://aws.amazon.com