Commit Graph
28 Commits
Author SHA1 Message Date
Bryan Endersstocker 07fedf66d0 Add NFC support to String.normalize/2 2015-10-03 11:08:55 -04:00
Bryan Endersstocker 945f8b6466 Optimize Unicode normalization
Also extracted String.Normalizer module from String.Unicode.
2015-09-26 10:59:04 -04:00
Bryan Endersstocker bf02047280 Add String.normalize/2 2015-09-26 00:12:44 -04:00
Bryan Enders dab2a632b6 Rename String.is_equivalent/2 to String.equivalent?/2 2015-09-25 15:01:29 -04:00
Bryan Enders f83cc7de5f Add String.is_equivalent/2 to test canonical Unicode equivalence 2015-09-25 13:35:16 -04:00
José Valim ff51926fe8 Move away from HashDict and HashSet 2015-09-25 12:46:22 +02:00
eksperimental a251ccc722 Format raise message consistently
Kernel.raise messages should not start with uppercase, and should not
have trailing punctuation.

Mix.raise messages should start with uppercase, and should not have a trailing
puntuation.

    # starting with uppercase
    # exclude (Mix.raise, and any message that starts with all uppecase such as: I or IO
    ag '(?<!Mix\.)raise\s+\"+(?!([A-Z]+\b))[A-Z]'

    # Mix.raise should start with upper case
    # except when `mix` or `rebar` are mentioned
    ag 'Mix\.raise\s+\"(?!(mix|rebar))+[a-z]'

    # ending in period
    ag "raise\s+\"+.+\.\"\s*\n"
2015-09-08 00:44:17 +07:00
José Valim 0e7232671b Optimize many functions in String
Many functions relied on next_grapheme, which would traverse
the whole string generating a lot of graphemes as garbage along
the way.

This commit introduces next_grapheme_size, which returns the
size in bytes of the next grapheme, avoiding the generation
of the garbage and optimizing functions that need just the
byte_size.

Functions optimized are: String.at, String.split_at, String.slice
and String.length.
2015-08-02 13:49:50 +02:00
José Valim a52bf5f40e Speed up upcase and downcase for large strings 2015-06-05 13:12:25 +02:00
José Valim 6efdf35f41 Handle corner cases for small strings in rstrp 2015-06-02 10:13:45 +02:00
José Valim 8574c7b568 Optimize rstrip
This new implementation is no longer linear without affecting
smaller samples. For a string that is 100 bytes long, it is
25x faster than the previous implementation.
2015-06-02 09:52:13 +02:00
eksperimental 50ab29ffb1 fix typo in previous commit. bodyless defmacro added to show proper arg in b/1, t/1 2015-03-28 19:19:37 +07:00
eksperimental 3592151b3e rename function arg. when "other" was used 2015-03-27 02:37:36 +07:00
jw2013 f2863de304 fix wrong CRLF grapheme 2014-11-08 12:45:41 -08:00
José Valim 9717a63604 Do not consider subpatterns on Regex.split/3 2014-07-30 11:13:18 +02:00
José Valim 182cbfe541 Update to Unicode 7.0.0 2014-06-17 17:17:32 +02:00
José Valim 1038018e84 Migrate from size to byte_size and tuple_size 2014-06-06 13:33:02 +02:00
José Valim 540c763c85 Move binary and list conversions to String and List 2014-05-17 17:02:46 +02:00
José Valim 3bb8e17e9e Add iodata_to_binary and chardata_to_string to IO 2014-05-17 13:06:40 +02:00
José Valim 56f15d9f4e Do not add spaces after { and before }
This makes the source code consistent with the result
returned by inspect/2.
2014-04-21 19:06:35 +02:00
José Valim 024edf3d40 Rename iolist_* to iodata_* 2014-02-23 11:32:03 +01:00
José Valim 3512c1a860 Start migrating lc to for comprehensions 2014-02-17 14:45:04 +01:00
José Valim ba54afbf45 Move to new sigils syntax 2014-02-05 14:39:04 +01:00
José Valim 0802014779 No longer compile unicode in parallel
Compiling in parallel only bought us 2s (out of 24s) since
compiling graphemes is what takes most of the time but it
considerably increases memory usage, by about 30%, since both
compilation targets are loaded into memory at once.
2014-01-22 17:51:24 +01:00
José Valim 87fee37711 Change String.next_grapheme/1 and String.next_codepoint/1 to return nil on string end
This is more composable in the sense can use in in if/cond constructs
and also makes it easy to build streams. Closes #1721.
2014-01-14 16:15:54 +01:00
José Valim 0538e46023 Compile unicode modules in parallel 2014-01-09 20:12:20 +01:00
José Valim a829da1775 Reduce the size of String.Graphemes module
Based on a suggestion @krestenkrab and an initial implementation
proposed in #1902, we can reduce the total size of the String.Graphemes
beam file by ensuring the function clauses for next_grapheme are
similar.

We have done this by implementing next_grapheme in a tail recursive
fashion and using binary_part to select the first bits. The end result
is a consistently 33% faster graphemes operation for ascii strings
and only 10% slower operation for heavy unicode strings (which may
be further optimazable).
2013-11-29 12:48:01 +01:00
José Valim 94579848cc Move unicode from priv/ to unicode/ 2013-11-05 08:17:33 +01:00