Received: from malur.postgresql.org ([217.196.149.56]) by arkaria.postgresql.org with esmtp (Exim 4.72) (envelope-from ) id 1UhLV4-00051I-9T for pgsql-sql@arkaria.postgresql.org; Tue, 28 May 2013 15:07:46 +0000 Received: from localhost ([127.0.0.1] helo=postgresql.org) by malur.postgresql.org with smtp (Exim 4.72) (envelope-from ) id 1UhLV3-0000en-Qh for pgsql-sql@arkaria.postgresql.org; Tue, 28 May 2013 15:07:45 +0000 Received: from makus.postgresql.org ([2001:4800:7903:4::125]) by malur.postgresql.org with esmtp (Exim 4.72) (envelope-from ) id 1UhLV2-0000eh-Oq for pgsql-sql@postgresql.org; Tue, 28 May 2013 15:07:44 +0000 Received: from mail-ea0-x22a.google.com ([2a00:1450:4013:c01::22a]) by makus.postgresql.org with esmtp (Exim 4.80) (envelope-from ) id 1UhLUs-00030e-8k for pgsql-sql@postgresql.org; Tue, 28 May 2013 15:07:44 +0000 Received: by mail-ea0-f170.google.com with SMTP id f15so4629164eak.15 for ; Tue, 28 May 2013 08:07:33 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20120113; h=from:to:subject:date:message-id:in-reply-to:references:mime-version :content-type:x-mailer; bh=N7MU4quE6k4stAET0pWvt9BgLsEohs/J6kZ4iBH8aVg=; b=teZvqsZv4DMIMZX47v3Bk6kpehgnGX8gWEaxQ3qPtiFoF2f50UtiZ+qb2TgUUx4yth uGDIme4g4tvC2hag7WdKfMveG20/caU3vjUWyuT3Uo1hjNi8rCjWzabkA7/M7GMvzfJ+ u4S2TY3ooHCnwfLJynzG/MgZ2oP4bKWPf+eYJUog2EXOtZvtZqiueoBM1Uzw2b48AY6v IRKyE1zOfyqi+mE1ZWjETF4hnbHd+pxQUGjSR07uJNj4txWxlSH+QuN/yaqrPR527dz0 D/oip3ZeFyyiXa2QeZOb3eiUgvd9xO6qG4LywB4G6mE9uVl5fHUZhjRByn1O6JEnNQ08 lBEg== X-Received: by 10.15.73.133 with SMTP id h5mr14450842eey.118.1369753653176; Tue, 28 May 2013 08:07:33 -0700 (PDT) Received: from [134.2.11.31] (closure.Informatik.Uni-Tuebingen.De. [134.2.11.31]) by mx.google.com with ESMTPSA id h49sm16241182eew.7.2013.05.28.08.07.32 for (version=TLSv1 cipher=RC4-SHA bits=128/128); Tue, 28 May 2013 08:07:32 -0700 (PDT) From: "Torsten Grust" To: pgsql-sql@postgresql.org Subject: Re: reduce many loosely related rows down to one Date: Tue, 28 May 2013 17:07:35 +0200 Message-ID: <5ED3E628-D3B4-44BE-992A-568B7FD42996@gmail.com> In-Reply-To: <51A0660A.2010806@dhs-club.com> References: <51A0660A.2010806@dhs-club.com> MIME-Version: 1.0 Content-Type: text/plain; format=flowed X-Mailer: MailMate (1.5.4r3361) X-Pg-Spam-Score: 0.7 (/) List-Archive: List-Help: List-ID: List-Owner: List-Post: List-Subscribe: List-Unsubscribe: X-Mailing-List: pgsql-sql Precedence: bulk Sender: pgsql-sql-owner@postgresql.org On 25 May 2013, at 9:19, Bill MacArthur wrote (with possible deletions): > [...] > select * from test; > > id | rspid | nspid | cid | iac | newp | oldp | ppv | tppv > ----+-------+-------+-----+-----+------+------+---------+--------- > 1 | 2 | 3 | 4 | t | | | | > 1 | 2 | 3 | | | 100 | | | > 1 | 2 | 3 | | | | 200 | | > 1 | 2 | 3 | | | | | | 4100.00 > 1 | 2 | 3 | | | | | | 3100.00 > 1 | 2 | 3 | | | | | -100.00 | > 1 | 2 | 3 | | | | | 250.00 | > 2 | 7 | 8 | 4 | | | | | > (8 rows) > > -- I want this result (where ppv and tppv are summed and the other > distinct values are boiled down into one row) > -- I want to avoid writing explicit UNIONs that will break if, say the > "cid" was entered as a discreet row from the row containing "iac" > -- in this example "rspid" and "nspid" are always the same for a given > ID, however they could possibly be absent for a given row as well > > id | rspid | nspid | cid | iac | newp | oldp | ppv | tppv > ----+-------+-------+-----+-----+------+------+---------+--------- > 1 | 2 | 3 | 4 | t | 100 | 200 | 150.00 | 7200.00 > 2 | 7 | 8 | 4 | | | | 0.00 | 0.00 One possible option could be SELECT id, (array_agg(rspid))[1] AS rspid, -- (1) (array_agg(nspid))[1] AS nspid, (array_agg(cid))[1] AS cid, bool_or(iac) AS iac, -- (2) max(newp) AS newp, -- (3) min(oldp) AS oldp, -- (4) coalesce(sum(ppv), 0) AS ppv, coalesce(sum(tppv),0) AS tppv FROM test GROUP BY id; This query computes the desired output for your example input. There's a caveat here: your description of the problem has been somewhat vague and it remains unclear how the query should respond if the functional dependency id -> rspid does not hold. In this case, the array_agg(rspid)[1] in the line marked (1) will pick one among many different(!) rspid values. I don't know your scenario well enough to judge whether this would be an acceptable behavior. Other possible behaviors have been implemented in the lines (2), (3), (4) where different aggregation functions are used to reduce sets to a single value (e.g., pick the largest/smallest of many values ...). Cheers, --Torsten -- | Torsten "Teggy" Grust | Torsten.Grust@gmail.com -- Sent via pgsql-sql mailing list (pgsql-sql@postgresql.org) To make changes to your subscription: http://www.postgresql.org/mailpref/pgsql-sql