José Valim
5a2b24d21f
Group Unicode upcase/downcase by prefix ( #11310 )
...
The patch computes byte lookups based on the prefix. For example,
Á, É, etc all have the same prefix <<195>>, so they are lumped
together for lookup and then we just do a byte lookup later. We
tried doing the byte lookup on 64-element tuple (since the byte
is always within 0b10000000 and 0b10111111) but that's slower,
especially because we need to check the byte range for invalid
Unicode, so instead the last byte lookup is a case. Grouping the
top-level lookup makes the cost of a miss 3x cheaper albeit a
hit is 10% more expensive and reduces bytecode size.
2021-10-14 14:44:13 +02:00
Eksperimental
d962ddb5b0
Upgrade Unicode database to v14.0.0 ( #11308 )
...
Closes #11288 .
2021-10-12 18:05:10 +02:00
José Valim
3f5b3f00f1
Depend on Erlang for computing grapheme clusters ( #11024 )
...
Erlang ships with its own embedding of the Unicode
Codebase for a couple releases and this release changes
Elixir to depend on it in order to compute grapheme
clusters.
The Erlang implementation was up to 2x faster in low
codepoints (such as latin1) while the Elixir one could
be faster up to 3x in high codepoints (such as emoji)
so at the end the performance results are roughly the
same. As a benefit, we no longer need to ship our copy
of the grapheme cluster algorithm, which would take up
to 250kB in disk and more than 15 seconds to compile.
Note we still keep our own String downcase and upcase
algorithms, as our version is considerably more efficient
on all cases since it works exclusively with binaries
(more than 5x faster).
2021-06-01 12:44:55 +02:00
José Valim
0123f8680f
Reduce size of Unicode module in 33% by not duplicating codepoints
2021-05-25 18:06:58 +02:00
Eksperimental
c1e622950c
Use proper Unicode reference ( #10924 )
2021-04-19 17:10:44 +02:00
Enrico Rivarola
218c35a041
Sync unicode data to Unicode 13.0.0 ( #10866 )
2021-04-03 15:20:05 +02:00
Eksperimental
0beb4f42fe
Update Unicode upgrade instructions ( #10862 )
2021-04-01 22:11:23 +02:00
Eksperimental
37a09feaaa
Update Unicode to version 13.0.0 ( #10861 )
2021-04-01 20:35:52 +02:00
Cẩm Huỳnh
ef16468175
Optimize String.codepoints ( #10633 )
2021-01-08 19:20:27 +01:00
José Valim
c36791091f
Consider CRLF on Windows, closes #10470
2020-11-01 17:25:06 +01:00
Eksperimental
b0d4288174
Use titlecase in Elixir, Erlang, EEx, Unicode, and Dialyzer ( #9557 )
2019-11-19 09:10:29 +01:00
José Valim
bffb79e553
Update to Unicode 12.1
2019-07-05 11:04:08 +02:00
José Valim
0e6b8732ba
Improve docs for built-in data types
2019-07-05 11:04:08 +02:00
José Valim
076dece778
Avoid unicode codepoint computation for lower bytes
2019-07-04 09:38:55 +02:00
Eric Meadows-Jönsson
9b6ec7bae7
Move unreachable function check from xref to group pass ( #9168 )
2019-06-28 10:25:49 +02:00
José Valim
93fb75d793
Optimize Unicode graphemes
...
We only check for Pictographic once we find a ZWJ.
2019-03-27 01:46:34 +01:00
Eksperimental
99d919e0c8
Use "code point" instead of "codepoint" ( #8774 )
...
That's the term it is used in the Unicode standard.
2019-02-06 11:21:03 +01:00
Xavier Noria
ee5a7fa0a4
Clarify String.next_codepoint/1 re invalid UTF-8 ( #8489 )
2018-12-08 16:41:57 +01:00
Sun Yaozhu
1967a06abe
Fix ZWJ handling in Unicode grapheme clusters ( #8361 )
2018-11-04 19:56:36 +01:00
José Valim
54cb02c240
Rely on Erlang/OTP normalization which is faster
2018-07-12 13:28:58 +02:00
José Valim
c995ba968f
Update to Unicode 11
2018-06-06 11:22:36 +02:00
Michał Muskała
722aef3791
Use binary_part/3 instead of :binary.part/3
...
binary_part is a guard BIF and is slightly more efficient to use then
:binary.part which does a fully qualified external call.
2018-04-30 13:22:19 +02:00
Guilherme Pasqualino
e08df95e86
Run code formatter on lib/elixir/unicode/unicode.ex ( #6795 )
2017-10-10 12:24:15 +02:00
Michał Muskała
c7a7a3625a
Add is_binary guard to String.length ( #6430 )
...
This produces a bit more readable message if not a string is passed into `String.length`, then failing in the `next_grapheme_size` function.
2017-08-03 21:04:12 +02:00
José Valim
a311ae403e
Update to Unicode 10 ( #6293 )
2017-07-04 09:57:07 +02:00
José Valim
9873e4239f
Add non-quoted Unicode atoms and variables ( #6158 )
...
It follows Unicode Annex #31 .
See the Unicode Syntax document for a more in-depth reference.
2017-05-27 22:43:05 +02:00
José Valim
7ca7139bb4
Properly slice titlecase_once binaries
2017-05-19 12:37:05 +02:00
José Valim
636c49fb5c
Add grapheme tests
2017-02-13 15:44:59 +01:00
José Valim
3e394e8ef0
Incorporate new grapheme rules in Unicode 9
2017-02-13 14:34:04 +01:00
Andrea Leopardi
3604554aea
Remove leftover trailing whitespace from the code base
...
[ci skip]
2017-01-13 15:01:47 +01:00
José Valim
bd7b4847bd
Speed up String.split/1
2016-11-21 19:03:16 +01:00
José Valim
1606c40453
Update to Unicode 9.0.0
...
The update instructions have also been outlined
on top of the unicode.ex file.
Closes #4864 .
2016-11-20 00:13:12 +01:00
Miles Starkenburg
551a9ac14e
Fix nfd normalization bug ( #4606 )
...
* Fix nfd normalization bug
* Change String.normalize return type to binary
2016-05-13 09:28:50 +02:00
eksperimental
6da34e0225
Standardize unicode and friends ( #4569 )
...
* Unicode is a proper noun, so it should be capitalized
* Replace 'char data' with 'chardata'
* Correct use of a/an articles
2016-05-05 09:21:05 +02:00
Aleksei Magusev
7e20b6ef99
Merge pull request #4535 from lexmag/string-trim-pad
...
Introduce String.pad_{leading,trailing}/3 and String.trim{,_leading,_trailing}/2
2016-04-25 20:02:19 +02:00
Aleksei Magusev
24ab310950
Introduce String.trim{,_leading,_trailing}/2
2016-04-25 19:59:24 +02:00
eksperimental
4ceb41e71b
Formmating: Add white space around vertical bar ( #4507 )
2016-04-25 00:55:49 +02:00
José Valim
8fe87ffa3c
Fold decomposition recursions at compile time
...
Signed-off-by: José Valim <jose.valim@plataformatec.com.br >
2016-03-29 16:55:29 +02:00
José Valim
79b132a665
Do not reorder starting classes
...
Signed-off-by: José Valim <jose.valim@plataformatec.com.br >
2016-03-29 13:38:14 +02:00
José Valim
cd953f5883
Handle compositios with non-zero combining class
...
Signed-off-by: José Valim <jose.valim@plataformatec.com.br >
2016-03-29 13:38:10 +02:00
José Valim
107a185209
More optimizations and improvements to normalization
...
Signed-off-by: José Valim <jose.valim@plataformatec.com.br >
2016-03-29 13:38:07 +02:00
José Valim
c07f2b2820
Make decomposition recursive and consider exclusion list
...
Signed-off-by: José Valim <jose.valim@plataformatec.com.br >
2016-03-29 13:38:00 +02:00
José Valim
ec027fa421
Use combining_classes from UnicodeData
...
Signed-off-by: José Valim <jose.valim@plataformatec.com.br >
2016-03-29 13:37:56 +02:00
José Valim
664feb1802
Rely only on UnicodeData for composition/decomposition
2016-03-29 00:54:20 +02:00
José Valim
8199b81f7c
Clean up unicode range parsing
2016-03-18 16:59:13 +01:00
José Valim
e1a1065ee9
Update Unicode to 8.0.0 (equivalent pending)
2016-03-15 15:58:01 +01:00
José Valim
646ae4d63f
Avoid regular expressions
2016-03-15 15:43:32 +01:00
Abel Muiño
98a4ee5da6
New definition of whitespace & breakable whitespace
...
All whitespace can be removed by `strip` but only breakable whitespace
can be used as a delimiter by `split`, with updated docs for String.split/1
2016-03-15 14:54:11 +01:00
José Valim
ea4102355c
Remove debug_info from unicode
2016-02-01 09:21:49 +01:00
José Valim
e65b01b870
Do not include debug_info in metadata String modules
2016-01-08 10:11:07 +01:00