Commit Graph
80 Commits
Author SHA1 Message Date
José Valim 5a2b24d21f Group Unicode upcase/downcase by prefix (#11310)
The patch computes byte lookups based on the prefix. For example,
Á, É, etc all have the same prefix <<195>>, so they are lumped
together for lookup and then we just do a byte lookup later. We
tried doing the byte lookup on 64-element tuple (since the byte
is always within 0b10000000 and 0b10111111) but that's slower,
especially because we need to check the byte range for invalid
Unicode, so instead the last byte lookup is a case. Grouping the
top-level lookup makes the cost of a miss 3x cheaper albeit a
hit is 10% more expensive and reduces bytecode size.
2021-10-14 14:44:13 +02:00
Eksperimental d962ddb5b0 Upgrade Unicode database to v14.0.0 (#11308)
Closes #11288.
2021-10-12 18:05:10 +02:00
José Valim 3f5b3f00f1 Depend on Erlang for computing grapheme clusters (#11024)
Erlang ships with its own embedding of the Unicode
Codebase for a couple releases and this release changes
Elixir to depend on it in order to compute grapheme
clusters.

The Erlang implementation was up to 2x faster in low
codepoints (such as latin1) while the Elixir one could
be faster up to 3x in high codepoints (such as emoji)
so at the end the performance results are roughly the
same. As a benefit, we no longer need to ship our copy
of the grapheme cluster algorithm, which would take up
to 250kB in disk and more than 15 seconds to compile.

Note we still keep our own String downcase and upcase
algorithms, as our version is considerably more efficient
on all cases since it works exclusively with binaries
(more than 5x faster).
2021-06-01 12:44:55 +02:00
José Valim 0123f8680f Reduce size of Unicode module in 33% by not duplicating codepoints 2021-05-25 18:06:58 +02:00
Eksperimental c1e622950c Use proper Unicode reference (#10924) 2021-04-19 17:10:44 +02:00
Enrico Rivarola 218c35a041 Sync unicode data to Unicode 13.0.0 (#10866) 2021-04-03 15:20:05 +02:00
Eksperimental 0beb4f42fe Update Unicode upgrade instructions (#10862) 2021-04-01 22:11:23 +02:00
Eksperimental 37a09feaaa Update Unicode to version 13.0.0 (#10861) 2021-04-01 20:35:52 +02:00
Cẩm Huỳnh ef16468175 Optimize String.codepoints (#10633) 2021-01-08 19:20:27 +01:00
José Valim c36791091f Consider CRLF on Windows, closes #10470 2020-11-01 17:25:06 +01:00
Eksperimental b0d4288174 Use titlecase in Elixir, Erlang, EEx, Unicode, and Dialyzer (#9557) 2019-11-19 09:10:29 +01:00
José Valim bffb79e553 Update to Unicode 12.1 2019-07-05 11:04:08 +02:00
José Valim 0e6b8732ba Improve docs for built-in data types 2019-07-05 11:04:08 +02:00
José Valim 076dece778 Avoid unicode codepoint computation for lower bytes 2019-07-04 09:38:55 +02:00
Eric Meadows-Jönsson 9b6ec7bae7 Move unreachable function check from xref to group pass (#9168) 2019-06-28 10:25:49 +02:00
José Valim 93fb75d793 Optimize Unicode graphemes
We only check for Pictographic once we find a ZWJ.
2019-03-27 01:46:34 +01:00
Eksperimental 99d919e0c8 Use "code point" instead of "codepoint" (#8774)
That's the term it is used in the Unicode standard.
2019-02-06 11:21:03 +01:00
Xavier Noria ee5a7fa0a4 Clarify String.next_codepoint/1 re invalid UTF-8 (#8489) 2018-12-08 16:41:57 +01:00
Sun Yaozhu 1967a06abe Fix ZWJ handling in Unicode grapheme clusters (#8361) 2018-11-04 19:56:36 +01:00
José Valim 54cb02c240 Rely on Erlang/OTP normalization which is faster 2018-07-12 13:28:58 +02:00
José Valim c995ba968f Update to Unicode 11 2018-06-06 11:22:36 +02:00
Michał Muskała 722aef3791 Use binary_part/3 instead of :binary.part/3
binary_part is a guard BIF and is slightly more efficient to use then
:binary.part which does a fully qualified external call.
2018-04-30 13:22:19 +02:00
Guilherme Pasqualino e08df95e86 Run code formatter on lib/elixir/unicode/unicode.ex (#6795) 2017-10-10 12:24:15 +02:00
Michał Muskała c7a7a3625a Add is_binary guard to String.length (#6430)
This produces a bit more readable message if not a string is passed into `String.length`, then failing in the `next_grapheme_size` function.
2017-08-03 21:04:12 +02:00
José Valim a311ae403e Update to Unicode 10 (#6293) 2017-07-04 09:57:07 +02:00
José Valim 9873e4239f Add non-quoted Unicode atoms and variables (#6158)
It follows Unicode Annex #31.

See the Unicode Syntax document for a more in-depth reference.
2017-05-27 22:43:05 +02:00
José Valim 7ca7139bb4 Properly slice titlecase_once binaries 2017-05-19 12:37:05 +02:00
José Valim 636c49fb5c Add grapheme tests 2017-02-13 15:44:59 +01:00
José Valim 3e394e8ef0 Incorporate new grapheme rules in Unicode 9 2017-02-13 14:34:04 +01:00
Andrea Leopardi 3604554aea Remove leftover trailing whitespace from the code base
[ci skip]
2017-01-13 15:01:47 +01:00
José Valim bd7b4847bd Speed up String.split/1 2016-11-21 19:03:16 +01:00
José Valim 1606c40453 Update to Unicode 9.0.0
The update instructions have also been outlined
on top of the unicode.ex file.

Closes #4864.
2016-11-20 00:13:12 +01:00
Miles Starkenburg 551a9ac14e Fix nfd normalization bug (#4606)
* Fix nfd normalization bug

* Change String.normalize return type to binary
2016-05-13 09:28:50 +02:00
eksperimental 6da34e0225 Standardize unicode and friends (#4569)
* Unicode is a proper noun, so it should be capitalized

* Replace 'char data' with 'chardata'

* Correct use of a/an articles
2016-05-05 09:21:05 +02:00
Aleksei Magusev 7e20b6ef99 Merge pull request #4535 from lexmag/string-trim-pad
Introduce String.pad_{leading,trailing}/3 and String.trim{,_leading,_trailing}/2
2016-04-25 20:02:19 +02:00
Aleksei Magusev 24ab310950 Introduce String.trim{,_leading,_trailing}/2 2016-04-25 19:59:24 +02:00
eksperimental 4ceb41e71b Formmating: Add white space around vertical bar (#4507) 2016-04-25 00:55:49 +02:00
José Valim 8fe87ffa3c Fold decomposition recursions at compile time
Signed-off-by: José Valim <jose.valim@plataformatec.com.br>
2016-03-29 16:55:29 +02:00
José Valim 79b132a665 Do not reorder starting classes
Signed-off-by: José Valim <jose.valim@plataformatec.com.br>
2016-03-29 13:38:14 +02:00
José Valim cd953f5883 Handle compositios with non-zero combining class
Signed-off-by: José Valim <jose.valim@plataformatec.com.br>
2016-03-29 13:38:10 +02:00
José Valim 107a185209 More optimizations and improvements to normalization
Signed-off-by: José Valim <jose.valim@plataformatec.com.br>
2016-03-29 13:38:07 +02:00
José Valim c07f2b2820 Make decomposition recursive and consider exclusion list
Signed-off-by: José Valim <jose.valim@plataformatec.com.br>
2016-03-29 13:38:00 +02:00
José Valim ec027fa421 Use combining_classes from UnicodeData
Signed-off-by: José Valim <jose.valim@plataformatec.com.br>
2016-03-29 13:37:56 +02:00
José Valim 664feb1802 Rely only on UnicodeData for composition/decomposition 2016-03-29 00:54:20 +02:00
José Valim 8199b81f7c Clean up unicode range parsing 2016-03-18 16:59:13 +01:00
José Valim e1a1065ee9 Update Unicode to 8.0.0 (equivalent pending) 2016-03-15 15:58:01 +01:00
José Valim 646ae4d63f Avoid regular expressions 2016-03-15 15:43:32 +01:00
Abel Muiño 98a4ee5da6 New definition of whitespace & breakable whitespace
All whitespace can be removed by `strip` but only breakable whitespace
can be used as a delimiter by `split`, with updated docs for String.split/1
2016-03-15 14:54:11 +01:00
José Valim ea4102355c Remove debug_info from unicode 2016-02-01 09:21:49 +01:00
José Valim e65b01b870 Do not include debug_info in metadata String modules 2016-01-08 10:11:07 +01:00