Commit Graph
119 Commits
Author SHA1 Message Date
EksperimentalandMaintenance App a6b68dd684 Update Unicode to version 15.1.0 (#12947)
This is an automated commit created by the Maintenance project
https://github.com/eksperimental/maintenance

Before merging, please read the release notes by visiting
<http://www.unicode.org/versions/Unicode15.1.0/>
and assess if additional changes are necessary in the code base.

Co-authored-by: Maintenance App <maintenance-beam@autistici.org>
2023-09-20 09:16:13 +02:00
José Valim 76a92b5a80 Deprecate String.capitalize/2 in favor of :string.titlecase/1
When capitalize/2 was written, Erlang did not provide titlecase
functions. capitalize/2 also downcases the rest of the string
while Erlang doesn't, which is a common contention point.

Given our titlecase implementation requires 40kb of additional
code in .beam files, it makes sense to unify it with Erlang's
(which we have already done for most non-critical functionality
in unicode).
2023-08-09 16:15:41 +02:00
Andrea Leopardi dd295f4bc7 Fix spelling of "hard-coded" 2023-07-09 16:08:12 +02:00
José Valim 7454333fd0 Require pin variable when accessing variable inside binary size in match, closes #12588 2023-05-26 16:35:36 +02:00
maintenance-beam-app e8d8e7d693 Update Unicode to version 15.0.0 (#12143)
This is an automated commit created by the Maintenance project
https://github.com/eksperimental/maintenance

Before merging, please read the release notes by visiting
<http://www.unicode.org/versions/Unicode15.0.0/>
and assess if additional changes are necessary in the code base.
2022-09-18 00:04:19 +02:00
sabiwara 29bf449f3f Evaluate formating charlists as ~c sigils (#12064) 2022-08-10 06:47:00 +02:00
José Valim fbf94613f3 Fix regression on identifier\\ being interpreted as a single identifier 2022-06-30 15:52:18 +02:00
José Valim 850c7ddfe4 Normalize unicode characters once in tokenizer 2022-05-23 12:14:39 +03:00
José Valim 25212f2db1 Do not attempt to suggest starting tokens
It is hard to suggest because it is not possible to
know if the user intended an operator, alias, or
identifier, so we can give incorrect feedback.
2022-05-23 12:14:39 +03:00
Luc Fueston 31cbdbd6e0 nfc and additional normalizations for identifiers (#11859)
* nfc by default
* suggest nfkc on unexpected tokens when nfkc would have worked
* support additional normalizations (just micro/mu for now)
* parser, formatter tests
* tweak Macro.inner_classify(atom) not to rely on not_nfc causing error
2022-05-23 09:44:44 +02:00
John Bampton 83467dec72 Remove trailing whitespace (#11782) 2022-04-25 15:25:31 +02:00
José Valim f067abaf11 Remove duplication in unicode tokenizer 2022-04-09 09:46:56 +02:00
José Valim bd5ee19108 Simplify process of copying SpecialCasing.txt 2022-02-13 09:01:05 +01:00
Eksperimental d806a6858f Improve Unicode module update instructions (#11622) 2022-02-13 08:54:24 +01:00
José Valim 2676f8b221 Allow only single- and highly restricted mixed-scripts in identifiers (UTS39 C3) (#11621) 2022-02-11 20:44:06 +01:00
José Valim bee0bb56be Do not emit warnings when formatting code
Closes #11595.
2022-01-26 13:37:53 +01:00
José Valim 851ac27272 Run unicode security linting on more tokens 2022-01-21 10:50:02 +01:00
Luc Fueston cba50b7a28 Warn on confusable non-ascii identifiers (UTS 39, C2) (#11582) 2022-01-21 09:37:25 +01:00
Luc Fueston 1b80c8c838 Don't allow restricted characters in identifiers (UTS 39, C1) (#11580) 2022-01-17 19:42:36 +01:00
José Valim a8b72c759a Group Unicode upcase/downcase by prefix (#11310)
The patch computes byte lookups based on the prefix. For example,
Á, É, etc all have the same prefix <<195>>, so they are lumped
together for lookup and then we just do a byte lookup later. We
tried doing the byte lookup on 64-element tuple (since the byte
is always within 0b10000000 and 0b10111111) but that's slower,
especially because we need to check the byte range for invalid
Unicode, so instead the last byte lookup is a case. Grouping the
top-level lookup makes the cost of a miss 3x cheaper albeit a
hit is 10% more expensive and reduces bytecode size.
2021-10-14 14:44:13 +02:00
Eksperimental 9b80f8a40b Upgrade Unicode database to v14.0.0 (#11308)
Closes #11288.
2021-10-12 18:05:10 +02:00
José Valim b2491fd839 Depend on Erlang for computing grapheme clusters (#11024)
Erlang ships with its own embedding of the Unicode
Codebase for a couple releases and this release changes
Elixir to depend on it in order to compute grapheme
clusters.

The Erlang implementation was up to 2x faster in low
codepoints (such as latin1) while the Elixir one could
be faster up to 3x in high codepoints (such as emoji)
so at the end the performance results are roughly the
same. As a benefit, we no longer need to ship our copy
of the grapheme cluster algorithm, which would take up
to 250kB in disk and more than 15 seconds to compile.

Note we still keep our own String downcase and upcase
algorithms, as our version is considerably more efficient
on all cases since it works exclusively with binaries
(more than 5x faster).
2021-06-01 12:44:55 +02:00
José Valim 5872d4947e Reduce size of Unicode module in 33% by not duplicating codepoints 2021-05-25 18:06:58 +02:00
José Valim 8cfaa07c41 Split on \r\n and \n accordingly during unicode compilation, closes #10991 2021-05-20 12:12:42 +02:00
Eksperimental d7cf95df0c Use proper Unicode reference (#10924) 2021-04-19 17:10:44 +02:00
Enrico Rivarola 589f9a6ce7 Sync unicode data to Unicode 13.0.0 (#10866) 2021-04-03 15:20:05 +02:00
Eksperimental 3ba7fe7aae Update Unicode upgrade instructions (#10862) 2021-04-01 22:11:23 +02:00
Eksperimental 92ec16c77a Update Unicode to version 13.0.0 (#10861) 2021-04-01 20:35:52 +02:00
Cẩm Huỳnh 5b59ce40da Optimize String.codepoints (#10633) 2021-01-08 19:20:27 +01:00
José Valim 0820b21a0d Consider CRLF on Windows, closes #10470 2020-11-01 17:25:06 +01:00
José Valim 255c09e452 Only check the mode after matching on the codepoint
This guarantees the compiler will not emit the code
that is first always checking the mode.
2020-10-03 10:20:08 +02:00
Alt D. Soy 7759bcccae Add :turkic mode option to String case functions (#10376)
This PR adds support of Turkic languages casing of the letter i according
to Unicode SpecialCasing.txt.
2020-10-03 09:52:18 +02:00
Eksperimental 0acab05f53 Use titlecase in Elixir, Erlang, EEx, Unicode, and Dialyzer (#9557) 2019-11-19 09:10:29 +01:00
José Valim f410cb13d1 Update to Unicode 12.1 2019-07-05 11:04:08 +02:00
José Valim 45ae6faeb5 Improve docs for built-in data types 2019-07-05 11:04:08 +02:00
José Valim 3f976942c2 Avoid unicode codepoint computation for lower bytes 2019-07-04 09:38:55 +02:00
Eric Meadows-Jönsson 1f974f869e Move unreachable function check from xref to group pass (#9168) 2019-06-28 10:25:49 +02:00
Eksperimental 3f55466eed Use https instead of http when servers are redirecting (#9094)
The servers are automatically redirecting from http:// to https://
2019-05-29 07:37:54 +02:00
José Valim 6a05455c2e Optimize Unicode graphemes
We only check for Pictographic once we find a ZWJ.
2019-03-27 01:46:34 +01:00
Eksperimental 266caa2f2f Use "code point" instead of "codepoint" (#8774)
That's the term it is used in the Unicode standard.
2019-02-06 11:21:03 +01:00
Xavier Noria b784c9315d Clarify String.next_codepoint/1 re invalid UTF-8 (#8489) 2018-12-08 16:41:57 +01:00
Sun Yaozhu 6760ba44c7 Fix ZWJ handling in Unicode grapheme clusters (#8361) 2018-11-04 19:56:36 +01:00
José Valim df26e1bcbd Print number of tests from graphemes tests 2018-10-02 12:25:50 +02:00
José Valim 99868597b1 Rely on Erlang/OTP normalization which is faster 2018-07-12 13:28:58 +02:00
José Valim c43136216b Update to Unicode 11 2018-06-06 11:22:36 +02:00
Michał Muskała 7cfb63c636 Use binary_part/3 instead of :binary.part/3
binary_part is a guard BIF and is slightly more efficient to use then
:binary.part which does a fully qualified external call.
2018-04-30 13:22:19 +02:00
José Valim 0db0b85f52 Consider case ignorable characters on Greek downcasing, closes #7149
Note this does not impact the runtime cost of other downcasing
operations. The final beam file grew only in 8kb.
2018-01-09 13:59:13 +01:00
José Valim 3f95b4a17d Improvements to greek handling of downcase, closes #7149 2017-12-26 23:11:17 +01:00
José Valim 8fd0141297 Optimize String.downcase/1
Thanks to @michalmuskala for feedback and review.
2017-12-20 12:58:17 +01:00
José Valim 2a6a545b8b Provide :greek and :ascii mappings in String upcase/downcase/capitalize
Closes #7105.
2017-12-20 11:54:06 +01:00