José Valim
bffb79e553
Update to Unicode 12.1
2019-07-05 11:04:08 +02:00
José Valim
0e6b8732ba
Improve docs for built-in data types
2019-07-05 11:04:08 +02:00
José Valim
076dece778
Avoid unicode codepoint computation for lower bytes
2019-07-04 09:38:55 +02:00
Eric Meadows-Jönsson
9b6ec7bae7
Move unreachable function check from xref to group pass ( #9168 )
2019-06-28 10:25:49 +02:00
José Valim
93fb75d793
Optimize Unicode graphemes
...
We only check for Pictographic once we find a ZWJ.
2019-03-27 01:46:34 +01:00
Eksperimental
99d919e0c8
Use "code point" instead of "codepoint" ( #8774 )
...
That's the term it is used in the Unicode standard.
2019-02-06 11:21:03 +01:00
Xavier Noria
ee5a7fa0a4
Clarify String.next_codepoint/1 re invalid UTF-8 ( #8489 )
2018-12-08 16:41:57 +01:00
Sun Yaozhu
1967a06abe
Fix ZWJ handling in Unicode grapheme clusters ( #8361 )
2018-11-04 19:56:36 +01:00
José Valim
54cb02c240
Rely on Erlang/OTP normalization which is faster
2018-07-12 13:28:58 +02:00
José Valim
c995ba968f
Update to Unicode 11
2018-06-06 11:22:36 +02:00
Michał Muskała
722aef3791
Use binary_part/3 instead of :binary.part/3
...
binary_part is a guard BIF and is slightly more efficient to use then
:binary.part which does a fully qualified external call.
2018-04-30 13:22:19 +02:00
Guilherme Pasqualino
e08df95e86
Run code formatter on lib/elixir/unicode/unicode.ex ( #6795 )
2017-10-10 12:24:15 +02:00
Michał Muskała
c7a7a3625a
Add is_binary guard to String.length ( #6430 )
...
This produces a bit more readable message if not a string is passed into `String.length`, then failing in the `next_grapheme_size` function.
2017-08-03 21:04:12 +02:00
José Valim
a311ae403e
Update to Unicode 10 ( #6293 )
2017-07-04 09:57:07 +02:00
José Valim
9873e4239f
Add non-quoted Unicode atoms and variables ( #6158 )
...
It follows Unicode Annex #31 .
See the Unicode Syntax document for a more in-depth reference.
2017-05-27 22:43:05 +02:00
José Valim
7ca7139bb4
Properly slice titlecase_once binaries
2017-05-19 12:37:05 +02:00
José Valim
636c49fb5c
Add grapheme tests
2017-02-13 15:44:59 +01:00
José Valim
3e394e8ef0
Incorporate new grapheme rules in Unicode 9
2017-02-13 14:34:04 +01:00
Andrea Leopardi
3604554aea
Remove leftover trailing whitespace from the code base
...
[ci skip]
2017-01-13 15:01:47 +01:00
José Valim
bd7b4847bd
Speed up String.split/1
2016-11-21 19:03:16 +01:00
José Valim
1606c40453
Update to Unicode 9.0.0
...
The update instructions have also been outlined
on top of the unicode.ex file.
Closes #4864 .
2016-11-20 00:13:12 +01:00
Miles Starkenburg
551a9ac14e
Fix nfd normalization bug ( #4606 )
...
* Fix nfd normalization bug
* Change String.normalize return type to binary
2016-05-13 09:28:50 +02:00
eksperimental
6da34e0225
Standardize unicode and friends ( #4569 )
...
* Unicode is a proper noun, so it should be capitalized
* Replace 'char data' with 'chardata'
* Correct use of a/an articles
2016-05-05 09:21:05 +02:00
Aleksei Magusev
7e20b6ef99
Merge pull request #4535 from lexmag/string-trim-pad
...
Introduce String.pad_{leading,trailing}/3 and String.trim{,_leading,_trailing}/2
2016-04-25 20:02:19 +02:00
Aleksei Magusev
24ab310950
Introduce String.trim{,_leading,_trailing}/2
2016-04-25 19:59:24 +02:00
eksperimental
4ceb41e71b
Formmating: Add white space around vertical bar ( #4507 )
2016-04-25 00:55:49 +02:00
José Valim
8fe87ffa3c
Fold decomposition recursions at compile time
...
Signed-off-by: José Valim <jose.valim@plataformatec.com.br >
2016-03-29 16:55:29 +02:00
José Valim
79b132a665
Do not reorder starting classes
...
Signed-off-by: José Valim <jose.valim@plataformatec.com.br >
2016-03-29 13:38:14 +02:00
José Valim
cd953f5883
Handle compositios with non-zero combining class
...
Signed-off-by: José Valim <jose.valim@plataformatec.com.br >
2016-03-29 13:38:10 +02:00
José Valim
107a185209
More optimizations and improvements to normalization
...
Signed-off-by: José Valim <jose.valim@plataformatec.com.br >
2016-03-29 13:38:07 +02:00
José Valim
c07f2b2820
Make decomposition recursive and consider exclusion list
...
Signed-off-by: José Valim <jose.valim@plataformatec.com.br >
2016-03-29 13:38:00 +02:00
José Valim
ec027fa421
Use combining_classes from UnicodeData
...
Signed-off-by: José Valim <jose.valim@plataformatec.com.br >
2016-03-29 13:37:56 +02:00
José Valim
664feb1802
Rely only on UnicodeData for composition/decomposition
2016-03-29 00:54:20 +02:00
José Valim
8199b81f7c
Clean up unicode range parsing
2016-03-18 16:59:13 +01:00
José Valim
e1a1065ee9
Update Unicode to 8.0.0 (equivalent pending)
2016-03-15 15:58:01 +01:00
José Valim
646ae4d63f
Avoid regular expressions
2016-03-15 15:43:32 +01:00
Abel Muiño
98a4ee5da6
New definition of whitespace & breakable whitespace
...
All whitespace can be removed by `strip` but only breakable whitespace
can be used as a delimiter by `split`, with updated docs for String.split/1
2016-03-15 14:54:11 +01:00
José Valim
ea4102355c
Remove debug_info from unicode
2016-02-01 09:21:49 +01:00
José Valim
e65b01b870
Do not include debug_info in metadata String modules
2016-01-08 10:11:07 +01:00
Aleksei Magusev
40096e0ff1
Eliminate spaces in bitstring definitions
2015-11-28 23:48:48 +01:00
José Valim
94a07df8d3
Remove uses of soon to be deprecated Dict
2015-10-23 15:04:24 +02:00
Bryan Endersstocker
07fedf66d0
Add NFC support to String.normalize/2
2015-10-03 11:08:55 -04:00
Bryan Endersstocker
945f8b6466
Optimize Unicode normalization
...
Also extracted String.Normalizer module from String.Unicode.
2015-09-26 10:59:04 -04:00
Bryan Endersstocker
bf02047280
Add String.normalize/2
2015-09-26 00:12:44 -04:00
Bryan Enders
dab2a632b6
Rename String.is_equivalent/2 to String.equivalent?/2
2015-09-25 15:01:29 -04:00
Bryan Enders
f83cc7de5f
Add String.is_equivalent/2 to test canonical Unicode equivalence
2015-09-25 13:35:16 -04:00
José Valim
ff51926fe8
Move away from HashDict and HashSet
2015-09-25 12:46:22 +02:00
eksperimental
a251ccc722
Format raise message consistently
...
Kernel.raise messages should not start with uppercase, and should not
have trailing punctuation.
Mix.raise messages should start with uppercase, and should not have a trailing
puntuation.
# starting with uppercase
# exclude (Mix.raise, and any message that starts with all uppecase such as: I or IO
ag '(?<!Mix\.)raise\s+\"+(?!([A-Z]+\b))[A-Z]'
# Mix.raise should start with upper case
# except when `mix` or `rebar` are mentioned
ag 'Mix\.raise\s+\"(?!(mix|rebar))+[a-z]'
# ending in period
ag "raise\s+\"+.+\.\"\s*\n"
2015-09-08 00:44:17 +07:00
José Valim
0e7232671b
Optimize many functions in String
...
Many functions relied on next_grapheme, which would traverse
the whole string generating a lot of graphemes as garbage along
the way.
This commit introduces next_grapheme_size, which returns the
size in bytes of the next grapheme, avoiding the generation
of the garbage and optimizing functions that need just the
byte_size.
Functions optimized are: String.at, String.split_at, String.slice
and String.length.
2015-08-02 13:49:50 +02:00
José Valim
a52bf5f40e
Speed up upcase and downcase for large strings
2015-06-05 13:12:25 +02:00