serenity

mirror of https://github.com/RGBCube/serenity synced 2025-10-30 06:52:36 +00:00

Author	SHA1	Message	Date
Timothy Flynn	a57615c2b4	Meta: Ensure cmake fails if we are unable to unzip the CLDR database	2021-08-26 23:40:23 +02:00
Timothy Flynn	ea21573ed8	LibUnicode: Download Unicode's CLDR database and generate locale data The Unicode standard publishes a database known as the Common Locale Data Repository (CLDR). This is a massive set of data from which anyone implementing Unicode's Technical Standard #35 may generate their implementation: https://www.unicode.org/reports/tr35/ This commit updates LibUnicode to download the compressed database and extract a small subset. That subset is used to generate a list of available locales and the territories (AKA regions) associated with each locale.	2021-08-26 22:04:09 +01:00
Timothy Flynn	a98d3a1a85	LibUnicode: Download and parse DerivedNormalizationProps UCD file This file contains the last properties that LibUnicode is not parsing. Much of the data in this file is not currently used; that is left as a FIXME for when String.prototype.normalize is implemented. Until then, only the code point properties are utilized for regular expression pattern escapes.	2021-08-11 13:11:01 +02:00
Timothy Flynn	7dce2bfe23	LibUnicode: Generate separate tables for General Category properties Previously, each code point's General Category was part of the generated UnicodeData structure. This ultimately presented two problems, one functional and one performance related: * Some General Categories are applied to unassigned code points, for example the Unassigned (Cn) category. Unassigned code points are strictly excluded from UnicodeData.txt, so by relying on that file, the generator is unable to handle these categories. * Lookups for General Categories are slower when searching through the large UnicodeData hash map. Even though lookups are O(1), the hash function turned out to be slower than binary searching through a category-specific table. So, now a table is generated for each General Category. When querying a code point for a category, a binary search is done on each code point range in that category's table to check if code point has that category. Further, General Categories are now parsed from the UCD file DerivedGeneralCategory.txt. This file is a normal "prop list" file and contains the categories for unassigned code points.	2021-08-11 13:11:01 +02:00
Timothy Flynn	4e546cee97	LibUnicode: Remove WordBreakProperty from generated Unicode data This was originally used for the "is_final_code_point" algorithm in LibUnicode/CharacterTypes.cpp. However, it has since been superseded by DerivedCoreProperties and is now unused. Remove it as it is currently a waste of time to process the data, and is trivial to add back if we need it again.	2021-08-11 13:11:01 +02:00
Timothy Flynn	6f2640d031	LibUnicode: Parse UCD DerivedBinaryProperties.txt and generate property	2021-08-04 13:50:32 +01:00
Timothy Flynn	9113f892a7	LibUnicode: Parse UCD emoji-data.txt and generate Unicode property	2021-08-04 13:50:32 +01:00
Timothy Flynn	5edd458420	LibUnicode: Parse UCD ScriptExtensions.txt and generate property	2021-08-04 13:50:32 +01:00
Timothy Flynn	f5c1bbc00b	LibUnicode: Parse UCD Scripts.txt and generate as a Unicode property There are a couple of minor nuances with parsing script values, compared to other properties. In Scripts.txt, the UCD file lists the full name of each script; other properties, like General Category, list the shorter name in their primary files. This means that the aliases listed in PropertyValueAliases.txt are reversed for script values.	2021-08-04 13:50:32 +01:00
Timothy Flynn	1bb6404a19	LibUnicode: Invoke Unicode data generator a single time It takes a non-neglible amount of time to parse all of the UCD files and generate the Unicode data files. To help compile times, only invoke the generator once.	2021-08-04 11:18:24 +02:00
Timothy Flynn	16e86ae743	LibUnicode: Generate General Category unions and aliases This downloads the PropertyValueAliases.txt UCD file, which contains a set of General Category aliases. This changes the General Category enumeration to now be generated as a bitmask. This is to easily allow General Category unions. For example, the LC (Cased_Letter) category is the union of the Ll, Lu, and Lt categories.	2021-08-02 21:02:09 +04:30
Timothy Flynn	bba3152104	LibUnicode: Parse and generate PropertyAliases These are all used for Unicode property escapes.	2021-07-30 21:26:31 +01:00
Timothy Flynn	761c16d873	LibUnicode: Parse and utilize DerivedCoreProperties DerivedCoreProperties are pseudo-properties that are the union of other categories and properties. For example, the derived property Math is the union of the general category Sm and the property Other_Math. Parsing these is necessary for implementing Unicode property escapes. But it also has the added benefit that LibUnicode now does not need to derive some of these properties at runtime.	2021-07-30 21:26:31 +01:00
Andrew Kaster	38707f4a20	LibUnicode: Make unicode data generation logic more relocatable The previous logic had several checks for Lagom directories and subdirectories. All we really want to do for these header checks is make sure that the files end up in an included folder prefixed with LibUnicode. We also don't need to hard code the path to the generator, the $<TARGET_FILES> generator expression can create the path for us.	2021-07-29 21:46:25 +01:00
Timothy Flynn	12fb3ae033	LibUnicode: Download and parse the word break property list UCD file Note that unlike the main property list, each code point has only one word break property. Code points that do not have a word break property are to be assigned the property "Other".	2021-07-28 23:42:29 +02:00
Timothy Flynn	38adfd8874	LibUnicode: Download and parse the property list UCD file	2021-07-28 23:42:29 +02:00
Timothy Flynn	32ea461385	LibUnicode: Download and parse the special casing UCD file This adds a SpecialCasing structure to the generated UnicodeData.h/cpp files. This structure contains casing rules for code points which have non-1-to-1 upper-to-lower case code point mappings. Further, these rules may be limited to specific locales or other context.	2021-07-27 21:04:36 +01:00
Timothy Flynn	98d8274040	Meta: Add LibUnicode (and its tests) to Lagom This is primarily to allow using LibUnicode within LibJS and its REPL. Note: this seems to be the first time that a Lagom dependency requires generated source files. For this to work, some of Lagom's CMakeLists.txt commands needed to be re-organized to include the CMake files that fetch and parse UnicodeData.txt. The paths required to invoke the generator also differ depending on what is currently building (SerenityOS vs. Lagom as part of the Serenity build vs. a standalone Lagom build).	2021-07-26 17:03:55 +01:00
Timothy Flynn	4dda3edc9e	LibUnicode: Introduce a Unicode library for interacting with UCD files The Unicode standard publishes the Unicode Character Database (UCD) with information about every code point, such as each code point's upper case mapping. LibUnicode exists to download and parse UCD files at build time and to provide accessors to that data. As a start, LibUnicode includes upper- and lower-case code point converters.	2021-07-26 17:03:55 +01:00

19 commits