pg.ddx.io pgsql-sql@postgresql.org mailing list archive
help / color / mirror / Atom feedUnderstanding Encoding
10+ messages / 6 participants
[nested] [flat]
* Understanding Encoding
@ 2013-09-06 05:56 Beena Emerson <memissemerson@gmail.com>
0 siblings, 3 replies; 10+ messages in thread
From: Beena Emerson @ 2013-09-06 05:56 UTC (permalink / raw)
To: pgsql-sql; pgsql-novice@postgresql.org
Hello All,
I am not able to understand how the encoding is handled. I would be happy
if someone can tell what is happening in the following scenario:
1. I have created a database with EUC_KR encoding and created a table and
inserted some korean value into it.
=# CREATE DATABASE korean WITH ENCODING 'EUC_KR' LC_COLLATE='ko_KR.euckr'
LC_CTYPE='ko_KR.euckr' TEMPLATE=template0;
=# \c korean
korean=# SHOW client_encoding;
client_encoding
-----------------
UTF8
(1 row)
korean=# CREATE TABLE tbl (doc text);
korean=# INSERT INTO tbl VALUES ('그레스');
2. If I insert non-korean values it throws error:
korean=# INSERT INTO tbl VALUES ('データベース');
ERROR: character with byte sequence 0xe3 0x83 0xbc in encoding "UTF8" has
no equivalent in encoding "EUC_KR"
korean=# SELECT * FROM tbl;
doc
--------
그레스
(1 row)
3. I change the client encoding to EUC_KR and try inserting the same korean
characters and it throws an error:
korean=# SET client_encoding = 'EUC_KR';
SET
korean=# INSERT INTO tbl VALUES ('그레스');
ERROR: invalid byte sequence for encoding "EUC_KR": 0xa0 0x88
Even the SELECT statement displays something different. I am not able to
understand why?
korean=# SELECT * FROM tbl;
doc
--------
����
(1 row)
Can someone please help me.
Thanks you,
Beena Emerson
^ permalink raw reply [nested|flat] 10+ messages in thread
* Re: Understanding Encoding
@ 2013-09-06 06:01 Gopal Tandon <gopal_tandon@ql2.com>
parent: Beena Emerson <memissemerson@gmail.com>
2 siblings, 0 replies; 10+ messages in thread
From: Gopal Tandon @ 2013-09-06 06:01 UTC (permalink / raw)
To: Beena Emerson <memissemerson@gmail.com>; +Cc: pgsql-sql; pgsql-novice@postgresql.org
You can refer : http://www.postgresql.org/docs/9.2/static/multibyte.html
On Fri, Sep 6, 2013 at 11:26 AM, Beena Emerson <memissemerson@gmail.com>wrote:
> Hello All,
>
> I am not able to understand how the encoding is handled. I would be happy
> if someone can tell what is happening in the following scenario:
>
> 1. I have created a database with EUC_KR encoding and created a table and
> inserted some korean value into it.
>
> =# CREATE DATABASE korean WITH ENCODING 'EUC_KR' LC_COLLATE='ko_KR.euckr'
> LC_CTYPE='ko_KR.euckr' TEMPLATE=template0;
>
> =# \c korean
>
> korean=# SHOW client_encoding;
> client_encoding
> -----------------
> UTF8
> (1 row)
>
> korean=# CREATE TABLE tbl (doc text);
>
> korean=# INSERT INTO tbl VALUES ('그레스');
>
>
> 2. If I insert non-korean values it throws error:
>
> korean=# INSERT INTO tbl VALUES ('データベース');
> ERROR: character with byte sequence 0xe3 0x83 0xbc in encoding "UTF8" has
> no equivalent in encoding "EUC_KR"
>
> korean=# SELECT * FROM tbl;
> doc
> --------
> 그레스
> (1 row)
>
>
> 3. I change the client encoding to EUC_KR and try inserting the same
> korean characters and it throws an error:
>
> korean=# SET client_encoding = 'EUC_KR';
> SET
> korean=# INSERT INTO tbl VALUES ('그레스');
> ERROR: invalid byte sequence for encoding "EUC_KR": 0xa0 0x88
>
>
> Even the SELECT statement displays something different. I am not able to
> understand why?
>
> korean=# SELECT * FROM tbl;
> doc
> --------
> ����
> (1 row)
>
>
> Can someone please help me.
>
> Thanks you,
>
> Beena Emerson
>
>
>
^ permalink raw reply [nested|flat] 10+ messages in thread
* Re: [NOVICE] Understanding Encoding
@ 2013-09-06 06:37 Amit Langote <amitlangote09@gmail.com>
parent: Beena Emerson <memissemerson@gmail.com>
2 siblings, 1 reply; 10+ messages in thread
From: Amit Langote @ 2013-09-06 06:37 UTC (permalink / raw)
To: Beena Emerson <memissemerson@gmail.com>; +Cc: pgsql-sql@postgresql.org, "pgsql-novice@postgresql.org" <pgsql-novice@postgresql.org>
On Fri, Sep 6, 2013 at 2:56 PM, Beena Emerson <memissemerson@gmail.com> wrote:
> Hello All,
>
> I am not able to understand how the encoding is handled. I would be happy if
> someone can tell what is happening in the following scenario:
>
> 1. I have created a database with EUC_KR encoding and created a table and
> inserted some korean value into it.
>
> =# CREATE DATABASE korean WITH ENCODING 'EUC_KR' LC_COLLATE='ko_KR.euckr'
> LC_CTYPE='ko_KR.euckr' TEMPLATE=template0;
>
> =# \c korean
>
> korean=# SHOW client_encoding;
> client_encoding
> -----------------
> UTF8
> (1 row)
>
> korean=# CREATE TABLE tbl (doc text);
>
> korean=# INSERT INTO tbl VALUES ('그레스');
>
>
> 2. If I insert non-korean values it throws error:
>
> korean=# INSERT INTO tbl VALUES ('データベース');
> ERROR: character with byte sequence 0xe3 0x83 0xbc in encoding "UTF8" has
> no equivalent in encoding "EUC_KR"
>
> korean=# SELECT * FROM tbl;
> doc
> --------
> 그레스
> (1 row)
>
>
> 3. I change the client encoding to EUC_KR and try inserting the same korean
> characters and it throws an error:
>
> korean=# SET client_encoding = 'EUC_KR';
> SET
> korean=# INSERT INTO tbl VALUES ('그레스');
> ERROR: invalid byte sequence for encoding "EUC_KR": 0xa0 0x88
>
>
> Even the SELECT statement displays something different. I am not able to
> understand why?
>
> korean=# SELECT * FROM tbl;
> doc
> --------
> ����
> (1 row)
>
I wonder if you have tried changing your "locale" to ko_KR; something like:
LANG=ko_KR LC_ALL=ko_KR \
psql -d korean
--
Amit Langote
--
Sent via pgsql-sql mailing list (pgsql-sql@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-sql
^ permalink raw reply [nested|flat] 10+ messages in thread
* Re: [NOVICE] Understanding Encoding
@ 2013-09-06 06:47 Beena Emerson <memissemerson@gmail.com>
parent: Amit Langote <amitlangote09@gmail.com>
0 siblings, 2 replies; 10+ messages in thread
From: Beena Emerson @ 2013-09-06 06:47 UTC (permalink / raw)
To: Amit Langote <amitlangote09@gmail.com>; +Cc: pgsql-sql@postgresql.org, "pgsql-novice@postgresql.org" <pgsql-novice@postgresql.org>
>
> I wonder if you have tried changing your "locale" to ko_KR; something like:
>
> LANG=ko_KR LC_ALL=ko_KR \
> psql -d korean
>
>
Hi,
It still gives same result:
$ LANG=ko_KR LC_ALL=ko_KR
$ psql -d korean
korean=# SHOW client_encoding;
client_encoding
-----------------
EUC_KR
(1 row)
korean=# INSERT INTO tbl VALUES ('그레스');
ERROR: invalid byte sequence for encoding "EUC_KR": 0xa0 0x88
Beena Emerson
^ permalink raw reply [nested|flat] 10+ messages in thread
* Re: [SQL] Understanding Encoding
@ 2013-09-06 06:59 Tom Lane <tgl@sss.pgh.pa.us>
parent: Beena Emerson <memissemerson@gmail.com>
1 sibling, 1 reply; 10+ messages in thread
From: Tom Lane @ 2013-09-06 06:59 UTC (permalink / raw)
To: Beena Emerson <memissemerson@gmail.com>; +Cc: Amit Langote <amitlangote09@gmail.com>; pgsql-sql@postgresql.org, "pgsql-novice@postgresql.org" <pgsql-novice@postgresql.org>
Beena Emerson <memissemerson@gmail.com> writes:
> It still gives same result:
> $ LANG=ko_KR LC_ALL=ko_KR
> $ psql -d korean
> korean=# SHOW client_encoding;
> client_encoding
> -----------------
> EUC_KR
> (1 row)
> korean=# INSERT INTO tbl VALUES ('그레스');
> ERROR: invalid byte sequence for encoding "EUC_KR": 0xa0 0x88
What you need to figure out is what encoding the text you are typing
is in. You're telling psql it's EUC_KR but it evidently isn't.
If you're typing these characters manually then it's probably determined
by a setting of the terminal-emulator program you're using. But if
you're copying-and-pasting then things get more complicated.
Also, what you did above is not what Amit suggested: he wanted you to put
the variable assignments on the same command line as the psql invocation,
so that they'd affect the environment passed to psql. I'm suspicious of
his solution because I'd have thought the terminal program would set up
the right environment ... but you might as well try it.
regards, tom lane
--
Sent via pgsql-novice mailing list (pgsql-novice@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-novice
^ permalink raw reply [nested|flat] 10+ messages in thread
* Re: Understanding Encoding
@ 2013-09-06 07:03 Tatsuo Ishii <ishii@postgresql.org>
parent: Beena Emerson <memissemerson@gmail.com>
2 siblings, 1 reply; 10+ messages in thread
From: Tatsuo Ishii @ 2013-09-06 07:03 UTC (permalink / raw)
To: memissemerson@gmail.com; +Cc: pgsql-sql; pgsql-novice@postgresql.org
> Hello All,
>
> I am not able to understand how the encoding is handled. I would be happy
> if someone can tell what is happening in the following scenario:
>
> 1. I have created a database with EUC_KR encoding and created a table and
> inserted some korean value into it.
>
> =# CREATE DATABASE korean WITH ENCODING 'EUC_KR' LC_COLLATE='ko_KR.euckr'
> LC_CTYPE='ko_KR.euckr' TEMPLATE=template0;
>
> =# \c korean
>
> korean=# SHOW client_encoding;
> client_encoding
> -----------------
> UTF8
> (1 row)
>
> korean=# CREATE TABLE tbl (doc text);
>
> korean=# INSERT INTO tbl VALUES ('그레스');
>
>
> 2. If I insert non-korean values it throws error:
>
> korean=# INSERT INTO tbl VALUES ('データベース');
> ERROR: character with byte sequence 0xe3 0x83 0xbc in encoding "UTF8" has
> no equivalent in encoding "EUC_KR"
The error messages says all. PostgreSQL accepted 'データベース'
encoded in UTF-8 then tried to convert to EUC_KR but failed, because
EUC_KR does not accept languages other than Korean (and ASCII). What
else did you expect?
> korean=# SELECT * FROM tbl;
> doc
> --------
> 그레스
> (1 row)
>
>
> 3. I change the client encoding to EUC_KR and try inserting the same korean
> characters and it throws an error:
>
> korean=# SET client_encoding = 'EUC_KR';
> SET
> korean=# INSERT INTO tbl VALUES ('그레스');
> ERROR: invalid byte sequence for encoding "EUC_KR": 0xa0 0x88
0xa0 is definitely not part of EUC_KR. That's why PostgreSQL throws an
error. I gues you are using UHC (Unified Hangul Code), rather than
EUC_KR. They are different encodings. You should do either:
1) Make sure that your termical encoding is EUC_KR.
2) set client_encoding = 'uhc';
> Even the SELECT statement displays something different. I am not able to
> understand why?
>
> korean=# SELECT * FROM tbl;
> doc
> --------
> ����
> (1 row)
This is because the same reason above.
> Can someone please help me.
>
> Thanks you,
>
> Beena Emerson
--
Tatsuo Ishii
SRA OSS, Inc. Japan
English: http://www.sraoss.co.jp/index_en.php
Japanese: http://www.sraoss.co.jp
--
Sent via pgsql-sql mailing list (pgsql-sql@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-sql
^ permalink raw reply [nested|flat] 10+ messages in thread
* Re: [NOVICE] Understanding Encoding
@ 2013-09-06 07:14 Amit Langote <amitlangote09@gmail.com>
parent: Beena Emerson <memissemerson@gmail.com>
1 sibling, 0 replies; 10+ messages in thread
From: Amit Langote @ 2013-09-06 07:14 UTC (permalink / raw)
To: Beena Emerson <memissemerson@gmail.com>; +Cc: pgsql-sql@postgresql.org, "pgsql-novice@postgresql.org" <pgsql-novice@postgresql.org>
On Fri, Sep 6, 2013 at 3:47 PM, Beena Emerson <memissemerson@gmail.com> wrote:
>
>>
>> I wonder if you have tried changing your "locale" to ko_KR; something
>> like:
>>
>> LANG=ko_KR LC_ALL=ko_KR \
>> psql -d korean
>>
>
> Hi,
>
> It still gives same result:
>
> $ LANG=ko_KR LC_ALL=ko_KR
> $ psql -d korean
>
> korean=# SHOW client_encoding;
> client_encoding
> -----------------
> EUC_KR
> (1 row)
>
> korean=# INSERT INTO tbl VALUES ('그레스');
> ERROR: invalid byte sequence for encoding "EUC_KR": 0xa0 0x88
I changed the encoding of the terminal emulator (GNOME Terminal
2.31.3) using the Terminal menu as:
Terminal -> Set Character Encoding -> Korean (EUC-KR)
Note that, if the menu only lists UTF-8, you'd have to add EUC-KR
using "Add or Remove".
And it seems to work; could you try the same?
--
Amit Langote
--
Sent via pgsql-sql mailing list (pgsql-sql@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-sql
^ permalink raw reply [nested|flat] 10+ messages in thread
* Re: [SQL] Understanding Encoding
@ 2013-09-06 07:24 Beena Emerson <memissemerson@gmail.com>
parent: Tom Lane <tgl@sss.pgh.pa.us>
0 siblings, 0 replies; 10+ messages in thread
From: Beena Emerson @ 2013-09-06 07:24 UTC (permalink / raw)
To: Tom Lane <tgl@sss.pgh.pa.us>; +Cc: Amit Langote <amitlangote09@gmail.com>; pgsql-sql@postgresql.org, "pgsql-novice@postgresql.org" <pgsql-novice@postgresql.org>
On Fri, Sep 6, 2013 at 12:29 PM, Tom Lane <tgl@sss.pgh.pa.us> wrote:
> Beena Emerson <memissemerson@gmail.com> writes:
> > It still gives same result:
>
> > $ LANG=ko_KR LC_ALL=ko_KR
> > $ psql -d korean
>
> > korean=# SHOW client_encoding;
> > client_encoding
> > -----------------
> > EUC_KR
> > (1 row)
>
> > korean=# INSERT INTO tbl VALUES ('그레스');
> > ERROR: invalid byte sequence for encoding "EUC_KR": 0xa0 0x88
>
> What you need to figure out is what encoding the text you are typing
> is in. You're telling psql it's EUC_KR but it evidently isn't.
> If you're typing these characters manually then it's probably determined
> by a setting of the terminal-emulator program you're using. But if
> you're copying-and-pasting then things get more complicated.
>
> Also, what you did above is not what Amit suggested: he wanted you to put
> the variable assignments on the same command line as the psql invocation,
> so that they'd affect the environment passed to psql. I'm suspicious of
> his solution because I'd have thought the terminal program would set up
> the right environment ... but you might as well try it.
>
I tried with both the assignment and invocation in same line. Again it gave
the same result.
Maybe the problem is with copy paste. I will look into it.
Thank you.
^ permalink raw reply [nested|flat] 10+ messages in thread
* Re: Understanding Encoding
@ 2013-09-06 07:53 Sebastien FLAESCH <sf@4js.com>
parent: Tatsuo Ishii <ishii@postgresql.org>
0 siblings, 1 reply; 10+ messages in thread
From: Sebastien FLAESCH @ 2013-09-06 07:53 UTC (permalink / raw)
To: pgsql-sql
Hi,
Tip:
To identify what encoding you enter in the psql command interpreter:
1) Open a file with vim
2) Type in you SQL or copy/paste
3) Save the file and quit vim
4) $ file <filename>
Should give you the encoding of that text file.
For ex:
sf@orca:~$ echo $LC_ALL
en_US.UTF-8
sf@orca:~$ cat /tmp/xx
abcdefé
sf@orca:~$ file /tmp/xx
/tmp/xx: UTF-8 Unicode text
Seb
On 09/06/2013 09:03 AM, Tatsuo Ishii wrote:
>> Hello All,
>>
>> I am not able to understand how the encoding is handled. I would be happy
>> if someone can tell what is happening in the following scenario:
>>
>> 1. I have created a database with EUC_KR encoding and created a table and
>> inserted some korean value into it.
>>
>> =# CREATE DATABASE korean WITH ENCODING 'EUC_KR' LC_COLLATE='ko_KR.euckr'
>> LC_CTYPE='ko_KR.euckr' TEMPLATE=template0;
>>
>> =# \c korean
>>
>> korean=# SHOW client_encoding;
>> client_encoding
>> -----------------
>> UTF8
>> (1 row)
>>
>> korean=# CREATE TABLE tbl (doc text);
>>
>> korean=# INSERT INTO tbl VALUES ('그레스');
>>
>>
>> 2. If I insert non-korean values it throws error:
>>
>> korean=# INSERT INTO tbl VALUES ('データベース');
>> ERROR: character with byte sequence 0xe3 0x83 0xbc in encoding "UTF8" has
>> no equivalent in encoding "EUC_KR"
>
> The error messages says all. PostgreSQL accepted 'データベース'
> encoded in UTF-8 then tried to convert to EUC_KR but failed, because
> EUC_KR does not accept languages other than Korean (and ASCII). What
> else did you expect?
>
>> korean=# SELECT * FROM tbl;
>> doc
>> --------
>> 그레스
>> (1 row)
>>
>>
>> 3. I change the client encoding to EUC_KR and try inserting the same korean
>> characters and it throws an error:
>>
>> korean=# SET client_encoding = 'EUC_KR';
>> SET
>> korean=# INSERT INTO tbl VALUES ('그레스');
>> ERROR: invalid byte sequence for encoding "EUC_KR": 0xa0 0x88
>
> 0xa0 is definitely not part of EUC_KR. That's why PostgreSQL throws an
> error. I gues you are using UHC (Unified Hangul Code), rather than
> EUC_KR. They are different encodings. You should do either:
>
> 1) Make sure that your termical encoding is EUC_KR.
>
> 2) set client_encoding = 'uhc';
>
>> Even the SELECT statement displays something different. I am not able to
>> understand why?
>>
>> korean=# SELECT * FROM tbl;
>> doc
>> --------
>> ����
>> (1 row)
>
> This is because the same reason above.
>
>> Can someone please help me.
>>
>> Thanks you,
>>
>> Beena Emerson
> --
> Tatsuo Ishii
> SRA OSS, Inc. Japan
> English: http://www.sraoss.co.jp/index_en.php
> Japanese: http://www.sraoss.co.jp
>
--
Sent via pgsql-sql mailing list (pgsql-sql@postgresql.org)
To make changes to your subscription:
http://www.postgresql.org/mailpref/pgsql-sql
^ permalink raw reply [nested|flat] 10+ messages in thread
* Re: Understanding Encoding
@ 2013-09-06 09:23 Beena Emerson <memissemerson@gmail.com>
parent: Sebastien FLAESCH <sf@4js.com>
0 siblings, 0 replies; 10+ messages in thread
From: Beena Emerson @ 2013-09-06 09:23 UTC (permalink / raw)
To: ; +Cc: pgsql-sql
Hello,
Thank you all.
Amit, Changing the encoding of the terminal emulator worked.
Sebastiean, the tip was helpful.
--
Beena Emerson
^ permalink raw reply [nested|flat] 10+ messages in thread
end of thread, other threads:[~2013-09-06 09:23 UTC | newest]
Thread overview: 10+ messages (download: mbox mbox.gz follow: Atom feed)
-- links below jump to the message on this page --
2013-09-06 05:56 Understanding Encoding Beena Emerson <memissemerson@gmail.com>
2013-09-06 06:01 ` Gopal Tandon <gopal_tandon@ql2.com>
2013-09-06 06:37 ` Amit Langote <amitlangote09@gmail.com>
2013-09-06 06:47 ` Beena Emerson <memissemerson@gmail.com>
2013-09-06 06:59 ` Tom Lane <tgl@sss.pgh.pa.us>
2013-09-06 07:24 ` Beena Emerson <memissemerson@gmail.com>
2013-09-06 07:14 ` Amit Langote <amitlangote09@gmail.com>
2013-09-06 07:03 ` Tatsuo Ishii <ishii@postgresql.org>
2013-09-06 07:53 ` Sebastien FLAESCH <sf@4js.com>
2013-09-06 09:23 ` Beena Emerson <memissemerson@gmail.com>
This inbox is served by DDX for PostgreSQL; see mirroring instructions
for how to clone and mirror all data and code used for this inbox