# neotoma-py Python bindings for the Neotoma parsing library. ## Installation ```bash pip install maturin maturin build --release target/wheels/neotoma_py-VERSION.whl ``` ## Usage ```python import neotoma # Parse a grammar and use it to parse input grammar = neotoma.Grammar('digits') result = grammar.parse("12345") print(result) # Or use the convenience function result = neotoma.parse_with_grammar('digits', "12345") print(result) ``` ## Grammar Syntax ### Terminals and Literals - `"terminal"` - matches literal string (with `\"` for escaped quotes) ### Built-in Character Classes - `digits` - matches one or more decimal digits (0-9) - `alpha` - matches one or more alphabetic characters (ASCII) - `alphanumeric` - matches one or more alphanumeric characters (ASCII) - `whitespace` - matches one or more whitespace characters (ASCII) - `hexdigits` - matches one or more hexadecimal digits (0-9, a-f, A-F) ### Unicode Character Classes - `udigits` - matches one or more decimal digits (Unicode) - `ualpha` - matches one or more alphabetic characters (Unicode) - `ualphanumeric` - matches one or more alphanumeric characters (Unicode) - `uwhitespace` - matches one or more whitespace characters (Unicode) ### Custom Character Classes - `[abc]` - matches any character in the set (custom character class) - `[^abc]` - matches any character NOT in the set (negated character class) ### Combinators - `(A B C)` - matches A followed by B followed by C (sequence) - `(| A B C)` - matches either A or B or C (alternatives) ### Repetition - `(* A)` - matches zero or more instances of A - `(+ A)` - matches one or more instances of A - `(? A)` - matches zero or one instances of A ### Separated Repetition - `(* A / B)` - matches zero or more A's separated by B (trailing B allowed but not required) - `(+ A / B)` - matches one or more A's separated by B (trailing B allowed but not required) ### Special Operators - `(< A)` - reads until condition A is met (UTF-8 aware) - `(< A B)` - reads until A, then parses the captured content with B ### Named Rules - `name = rule` - Names a parsing rule, for use in other rules and/or recursive rule definitions - `name` - References a named rule. The built-in keywords (`digits`, `alpha`, etc.) are not valid rule names - `@start name` - Identifies the starting rule for the grammar. If not provided, defaults to the last rule defined ### Examples ```python import neotoma # Simple digit parsing grammar = neotoma.Grammar('digits') result = grammar.parse("12345") # Custom grammar with rules grammar_text = ''' number = digits word = alpha sentence = (+ word whitespace) @start sentence ''' grammar = neotoma.Grammar(grammar_text) result = grammar.parse("hello world test") # Character class example grammar = neotoma.Grammar('[aeiou]') # vowels only result = grammar.parse("a") # Sequence and alternatives grammar = neotoma.Grammar('(| "hello" "hi" "hey")') result = grammar.parse("hello") # Repetition with separator grammar = neotoma.Grammar('(+ digits / ",")') # comma-separated numbers result = grammar.parse("1,2,3,4") ```