- If you want to handle Unicode and avoid The Unicode Bug, in which your strings sometimes act like they aren't actually Unicode: in perl 5.12+,
use feature 'unicode_strings';. For older perl, see Unicode::Semantics, or use utf8::upgrade by hand. These methods achieve their task by forcing "the UTF-8 flag" on for the string. - If you want strings in your source text with non-ASCII: save it as a utf-8 encoded file and
use utf8;. Or you can encode Unicode code points with hex-escapes,\xae→ ®, or\x{30ab}→ カ. There are technically other options, which have additional drawbacks (utf-16 breaks the #! line; latin-1 is restricted to latin-1 unless you decode it yourself.) - If you want to print to a UTF-8 aware environment like your terminal emulator or CGI STDOUT after issuing a
Content-Type: text/html; charset=utf-8header: setting UTF-8 on the filehandle withbinmode(STDOUT, ':utf8')is the minimum, but:encoding(utf-8)instead of:utf8makes stricter guarantees that real code points are coming out. - If you want to read a UTF-16 encoded document into a Unicode string with minimal fuss:
open(FH, '< :encoding(utf-16)', $name). Note that the document has to be correctly encoded. You can use the Encode module'sdecodefunction if you need finer control over error behavior, but that's naturally more fuss:use Encode; open(FH, '<', $name); while (<fh>) { $line = decode($_, 'utf-16', $POLICY); ... } - If you want to convert a Unicode string to a specific set of bytes for some encoding-unaware module to throw on the wire, use the encode function from the Encode module:
use Encode; $message->attr('content-type.charset', 'utf-16'); $message->data(encode("UTF-16", $body));(This example would be for MIME::Lite, if you're curious.) - If you want to read a file encoded with charset X, into a string encoded with charset Y, I've found no instant way to do this. It's probably best to pass the input-encoding along as the output-encoding if at all possible. But you might find the Encode module's
from_to(), or string-IO as inIO::File->new(\$out, '>:'), or maybe a whole PerlIO filter as in PerlIO::code helpful if you can't. - If you see "Wide character in ..." warnings, then you passed a string with code points >=0x100 to something that expected a byte string of some sort: either really latin-1, or an encoded string.
- If you see longer strings of gibberish where you expected sensible non-ASCII characters, then you have probably double-encoded, either literally, or by printing an encoded string to a filehandle which does encoding.
- If you see the Unicode replacement character in a stream that should be UTF-8, you haven't encoded at all, such as printing a byte string on a raw filehandle in an environment expecting UTF-8. Most likely, the filehandle should have an encoding set on it, per point #3 above, though that may cause #8 on other strings you've printed.
- If you are using modules, they each may or may not deal with Unicode. DBD::mysql has the
mysql_enable_utf8option; Email::MIME accepts encoded strings via body, and decoded ones through body_str, but for the latter, you must also set the charset and encoding attributes (which correspond to the charset of Content-Type, and the Content-Transfer-Encoding, respectively.) MIME::Lite does not handle decoded strings at all and hopes for the best.
Showing posts with label character sets. Show all posts
Showing posts with label character sets. Show all posts
Wednesday, January 18, 2012
Perl and Unicode in Brief
Perl requires a knob for every I/O, and expects you to set them all correctly yourself. By default, they're all off (Unicode-unaware) for backwards compatibility.
Tuesday, October 25, 2011
Character Sets: Get PHP, Perl, MySQL, and Unicode to Play Together
This post is a companion to Perl and Unicode in Brief, an attempt to cover the same ground more concisely.
This is an extended remix of my recent post on the subject, only less of a rambling story and more focused. Again, I'll start with some background definitions.
I'll also assume that you're going to make everything UTF-8, because as a US-centric American who has the luxury of using English, that's what makes the most sense for my systems. However, if you understand everything I wrote, it should not be difficult to make everything UTF-16 or any other encoding you desire.
This is an extended remix of my recent post on the subject, only less of a rambling story and more focused. Again, I'll start with some background definitions.
I'll also assume that you're going to make everything UTF-8, because as a US-centric American who has the luxury of using English, that's what makes the most sense for my systems. However, if you understand everything I wrote, it should not be difficult to make everything UTF-16 or any other encoding you desire.
Tuesday, October 18, 2011
Character Sets, Encodings, MySQL, and your data
This post is a companion to Perl and Unicode in Brief, an attempt to cover similar ground more concisely. And this post is a revised version of the one you're currently reading.
I'm currently moving data from a (relatively old now) MySQL 5.0 server into Amazon RDS. I've been here before, when I was moving data from MySQL 4.x into 5.0 and mangling character sets. This time, I want to make 100% sure everything comes across with maximum fidelity, and also get the character encoding as stored to be labeled correctly in MySQL.
First, a quick definition or two:
I'm currently moving data from a (relatively old now) MySQL 5.0 server into Amazon RDS. I've been here before, when I was moving data from MySQL 4.x into 5.0 and mangling character sets. This time, I want to make 100% sure everything comes across with maximum fidelity, and also get the character encoding as stored to be labeled correctly in MySQL.
First, a quick definition or two:
- Character Set: a specific table to translate between characters and numbers. Example: ASCII defines characters for numbers 0-127; "A" is 65. This can also be described as "a set of characters, and their corresponding representation inside the computer."
- Character Encoding: a means of "packing" numbers from the character set into a container. Example: UTF-8. The Unicode character 0x2013 becomes 0xE2,80,99. The "E" signifies "Part 1 of 3", and part of the remaining bytes simply indicate "Continued"; the 0x2013 is then divided up to fit in the parts of the bytes that aren't indicating their "Part 1" or "Continued" status. In the specific case of UTF-8, the encoding is designed so that the ASCII range 0-127 (0x00-7F) is encoded without change: a leading 0-7 means "Part 1 of 1".
- 8-bit character encoding: In older, simpler days, character sets defined only as many characters as could fit in 8 bits, and defined the encoding as simply the numbers. Character number 181 would encode as a byte (8 bits) with value 181.
- A character encoding implies the associated character set, because the encoding defines how numbers in its character set become individual bytes. How characters in other sets would be encoded is left undefined and basically impossible.
Subscribe to:
Posts (Atom)