## New Parser in Python

From: "andrew cooke" <andrew@...>

Date: Mon, 12 Jan 2009 00:54:55 -0300 (CLST)

Not ready for release yet, but I've just got a new parser, written in
Python, to the point where it's useful.

It includes full backtracing (parse forests etc), the (untested) ability
to 'automatically' control resource use (think 'maximum backtrace stack')
and enough syntactic sugar to rot your teeth :o)

This test:

from logging import basicConfig, DEBUG
from unittest import TestCase

from lepl.match import *
from lepl.node import Node

class NodeTest(TestCase):

def test_node(self):
basicConfig(level=DEBUG)

class Term(Node): pass
class Factor(Node): pass
class Expression(Node): pass

expression  = Delayed()
number      = Digit()[1:,...]                   > 'number'
term        = (number | '(' / expression / ')') > Term
muldiv      = Any('*/')                         > 'operator'
factor      = (term / (muldiv / term)[0:])      > Factor
expression += (factor / (addsub / factor)[0:])  > Expression

(ast, _) = next(expression.match_string('1 + 2 * (3 + 4 - 5)'))
print(ast[0])
(ast, _) = next(expression('1 + 2 * (3 + 4 - 5)'))
print(ast[0])

Prints the following (twice):

Expression
+- Factor
|   +- Term
|   |   - number=1
|   - ' '
+- operator=+
+- ' '
- Factor
+- Term
|   - number=2
+- ' '
+- operator=*
+- ' '
- Term
+- '('
+- Expression
|   +- Factor
|   |   +- Term
|   |   |   - number=3
|   |   - ' '
|   +- operator=+
|   +- ' '
|   +- Factor
|   |   +- Term
|   |   |   - number=4
|   |   - ' '
|   +- operator=-
|   +- ' '
|   - Factor
|       - Term
|           - number=5
- ')'

Andrew

### Parsing Credits

From: "andrew cooke" <andrew@...>

Date: Mon, 12 Jan 2009 00:58:24 -0300 (CLST)

I should have added that this copies lots of good ideas from both
pyparsing - http://pyparsing.wikispaces.com/ - and (more so) "Pattern
Matching In Python" http://www.wilmott.ca/python/patternmatching.html

Andrew

### Syntax

From: "andrew cooke" <andrew@...>

Date: Mon, 12 Jan 2009 01:11:39 -0300 (CLST)

A quick explanation of the syntax:

This allows forward references to 'expression', which will be defined later.

expression  = Delayed()

This defines 'number' as one or more digits, specified via '[1:]', and
combines the digits into a single string, specified via '[...]'.  The
result is then associated with the tag 'number'.

number      = Digit()[1:,...]                   > 'number'

This defines term as either 'number' or (with backtracing) a bracketed
expression.  The strings are automatically promoted to literal matches and
the '/' indicate that there are optional spaces between the matchers ('//'
for required space).  The result is used to construct a Term instance,
which is a subclass of Node (and which will automatically construct
attributes for the contents).

term        = (number | '(' / expression / ')') > Term

Define 'muldiv' to be either '*' or '/' and tag the result.

muldiv      = Any('*/')                         > 'operator'

Hopefully this is becoming obvious.  The '[0:]' here means '0 or more'
instances of 'muldiv', optional space, and 'term'.

factor      = (term / (muldiv / term)[0:])      > Factor

Nothing new here.

This defines the 'Delayed' matcher introduced earlier (it was introduced
so that we could reference it in 'term', even though we cannot define it
until later).

expression += (factor / (addsub / factor)[0:])  > Expression

Not sure if it's obvious, but one major aim has been to try to combine the
best of both OO and functional programming, in what I feel is a very
'Pythonic' way.

Andrew

### With Bactracking

From: "andrew cooke" <andrew@...>

Date: Mon, 12 Jan 2009 01:25:59 -0300 (CLST)

Changing the spec slightly to:

expression  = Delayed()
number      = Digit()[1:,...]                   > 'number'
term        = (number | '(' / expression / ')') > Term
muldiv      = Any('*/')                         > 'operator'
factor      = (term / (muldiv / term)[0:])      > Factor
expression += Drop(Any()[0:]) & \
(factor / (addsub / factor)[0:])  > Expression

And using:

for (ast, _) in expression('1 + 2 * (3 + 4 - 5)'):
print(ast[0])

Gives:

Expression
- Factor
- Term
- number '5'
Expression
+- Factor
|   +- Term
|   |   - number '4'
|   - ' '
+- operator '-'
+- ' '
- Factor
- Term
- number '5'
Expression
- Factor
+- Term
|   - number '4'
- ' '
Expression
+- Factor
|   - Term
|       - number '4'
+- ' '
+- operator '-'
+- ' '
- Factor
- Term
- number '5'
Expression
+- Factor
|   - Term
|       - number '4'
- ' '
Expression
- Factor
- Term
- number '4'
Expression
+- Factor
|   +- Term
|   |   - number '3'
|   - ' '
+- operator '+'
+- ' '
+- Factor
|   +- Term
|   |   - number '4'
|   - ' '
+- operator '-'
+- ' '
- Factor
- Term
- number '5'
Expression
+- Factor
|   +- Term
|   |   - number '3'
|   - ' '
+- operator '+'
+- ' '
- Factor
+- Term
|   - number '4'
- ' '
Expression
+- Factor
|   +- Term
|   |   - number '3'
|   - ' '
+- operator '+'
+- ' '
- Factor
- Term
- number '4'
Expression
- Factor
+- Term
|   - number '3'
- ' '
Expression
+- Factor
|   - Term
|       - number '3'
+- ' '
+- operator '+'
+- ' '
+- Factor
|   +- Term
|   |   - number '4'
|   - ' '
+- operator '-'
+- ' '
- Factor
- Term
- number '5'
Expression
+- Factor
|   - Term
|       - number '3'
+- ' '
+- operator '+'
+- ' '
- Factor
+- Term
|   - number '4'
- ' '
Expression
+- Factor
|   - Term
|       - number '3'
+- ' '
+- operator '+'
+- ' '
- Factor
- Term
- number '4'
Expression
+- Factor
|   - Term
|       - number '3'
- ' '
Expression
- Factor
- Term
- number '3'
Expression
- Factor
- Term
+- '('
+- Expression
|   - Factor
|       - Term
|           - number '5'
- ')'
Expression
- Factor
- Term
+- '('
+- Expression
|   +- Factor
|   |   +- Term
|   |   |   - number '4'
|   |   - ' '
|   +- operator '-'
|   +- ' '
|   - Factor
|       - Term
|           - number '5'
- ')'
Expression
- Factor
- Term
+- '('
+- Expression
|   +- Factor
|   |   - Term
|   |       - number '4'
|   +- ' '
|   +- operator '-'
|   +- ' '
|   - Factor
|       - Term
|           - number '5'
- ')'
Expression
- Factor
- Term
+- '('
+- Expression
|   +- Factor
|   |   +- Term
|   |   |   - number '3'
|   |   - ' '
|   +- operator '+'
|   +- ' '
|   +- Factor
|   |   +- Term
|   |   |   - number '4'
|   |   - ' '
|   +- operator '-'
|   +- ' '
|   - Factor
|       - Term
|           - number '5'
- ')'
Expression
- Factor
- Term
+- '('
+- Expression
|   +- Factor
|   |   - Term
|   |       - number '3'
|   +- ' '
|   +- operator '+'
|   +- ' '
|   +- Factor
|   |   +- Term
|   |   |   - number '4'
|   |   - ' '
|   +- operator '-'
|   +- ' '
|   - Factor
|       - Term
|           - number '5'
- ')'
Expression
- Factor
+- Term
|   - number '2'
+- ' '
+- operator '*'
+- ' '
- Term
+- '('
+- Expression
|   - Factor
|       - Term
|           - number '5'
- ')'
Expression
- Factor
+- Term
|   - number '2'
+- ' '
+- operator '*'
+- ' '
- Term
+- '('
+- Expression
|   +- Factor
|   |   +- Term
|   |   |   - number '4'
|   |   - ' '
|   +- operator '-'
|   +- ' '
|   - Factor
|       - Term
|           - number '5'
- ')'
Expression
- Factor
+- Term
|   - number '2'
+- ' '
+- operator '*'
+- ' '
- Term
+- '('
+- Expression
|   +- Factor
|   |   - Term
|   |       - number '4'
|   +- ' '
|   +- operator '-'
|   +- ' '
|   - Factor
|       - Term
|           - number '5'
- ')'
Expression
- Factor
+- Term
|   - number '2'
+- ' '
+- operator '*'
+- ' '
- Term
+- '('
+- Expression
|   +- Factor
|   |   +- Term
|   |   |   - number '3'
|   |   - ' '
|   +- operator '+'
|   +- ' '
|   +- Factor
|   |   +- Term
|   |   |   - number '4'
|   |   - ' '
|   +- operator '-'
|   +- ' '
|   - Factor
|       - Term
|           - number '5'
- ')'
Expression
- Factor
+- Term
|   - number '2'
+- ' '
+- operator '*'
+- ' '
- Term
+- '('
+- Expression
|   +- Factor
|   |   - Term
|   |       - number '3'
|   +- ' '
|   +- operator '+'
|   +- ' '
|   +- Factor
|   |   +- Term
|   |   |   - number '4'
|   |   - ' '
|   +- operator '-'
|   +- ' '
|   - Factor
|       - Term
|           - number '5'
- ')'
Expression
- Factor
+- Term
|   - number '2'
- ' '
Expression
+- Factor
|   - Term
|       - number '2'
- ' '
Expression
- Factor
- Term
- number '2'
Expression
+- Factor
|   +- Term
|   |   - number '1'
|   - ' '
+- operator '+'
+- ' '
- Factor
+- Term
|   - number '2'
+- ' '
+- operator '*'
+- ' '
- Term
+- '('
+- Expression
|   - Factor
|       - Term
|           - number '5'
- ')'
Expression
+- Factor
|   +- Term
|   |   - number '1'
|   - ' '
+- operator '+'
+- ' '
- Factor
+- Term
|   - number '2'
+- ' '
+- operator '*'
+- ' '
- Term
+- '('
+- Expression
|   +- Factor
|   |   +- Term
|   |   |   - number '4'
|   |   - ' '
|   +- operator '-'
|   +- ' '
|   - Factor
|       - Term
|           - number '5'
- ')'
Expression
+- Factor
|   +- Term
|   |   - number '1'
|   - ' '
+- operator '+'
+- ' '
- Factor
+- Term
|   - number '2'
+- ' '
+- operator '*'
+- ' '
- Term
+- '('
+- Expression
|   +- Factor
|   |   - Term
|   |       - number '4'
|   +- ' '
|   +- operator '-'
|   +- ' '
|   - Factor
|       - Term
|           - number '5'
- ')'
Expression
+- Factor
|   +- Term
|   |   - number '1'
|   - ' '
+- operator '+'
+- ' '
- Factor
+- Term
|   - number '2'
+- ' '
+- operator '*'
+- ' '
- Term
+- '('
+- Expression
|   +- Factor
|   |   +- Term
|   |   |   - number '3'
|   |   - ' '
|   +- operator '+'
|   +- ' '
|   +- Factor
|   |   +- Term
|   |   |   - number '4'
|   |   - ' '
|   +- operator '-'
|   +- ' '
|   - Factor
|       - Term
|           - number '5'
- ')'
Expression
+- Factor
|   +- Term
|   |   - number '1'
|   - ' '
+- operator '+'
+- ' '
- Factor
+- Term
|   - number '2'
+- ' '
+- operator '*'
+- ' '
- Term
+- '('
+- Expression
|   +- Factor
|   |   - Term
|   |       - number '3'
|   +- ' '
|   +- operator '+'
|   +- ' '
|   +- Factor
|   |   +- Term
|   |   |   - number '4'
|   |   - ' '
|   +- operator '-'
|   +- ' '
|   - Factor
|       - Term
|           - number '5'
- ')'
Expression
+- Factor
|   +- Term
|   |   - number '1'
|   - ' '
+- operator '+'
+- ' '
- Factor
+- Term
|   - number '2'
- ' '
Expression
+- Factor
|   +- Term
|   |   - number '1'
|   - ' '
+- operator '+'
+- ' '
- Factor
- Term
- number '2'
Expression
- Factor
+- Term
|   - number '1'
- ' '
Expression
+- Factor
|   - Term
|       - number '1'
+- ' '
+- operator '+'
+- ' '
- Factor
+- Term
|   - number '2'
+- ' '
+- operator '*'
+- ' '
- Term
+- '('
+- Expression
|   - Factor
|       - Term
|           - number '5'
- ')'
Expression
+- Factor
|   - Term
|       - number '1'
+- ' '
+- operator '+'
+- ' '
- Factor
+- Term
|   - number '2'
+- ' '
+- operator '*'
+- ' '
- Term
+- '('
+- Expression
|   +- Factor
|   |   +- Term
|   |   |   - number '4'
|   |   - ' '
|   +- operator '-'
|   +- ' '
|   - Factor
|       - Term
|           - number '5'
- ')'
Expression
+- Factor
|   - Term
|       - number '1'
+- ' '
+- operator '+'
+- ' '
- Factor
+- Term
|   - number '2'
+- ' '
+- operator '*'
+- ' '
- Term
+- '('
+- Expression
|   +- Factor
|   |   - Term
|   |       - number '4'
|   +- ' '
|   +- operator '-'
|   +- ' '
|   - Factor
|       - Term
|           - number '5'
- ')'
Expression
+- Factor
|   - Term
|       - number '1'
+- ' '
+- operator '+'
+- ' '
- Factor
+- Term
|   - number '2'
+- ' '
+- operator '*'
+- ' '
- Term
+- '('
+- Expression
|   +- Factor
|   |   +- Term
|   |   |   - number '3'
|   |   - ' '
|   +- operator '+'
|   +- ' '
|   +- Factor
|   |   +- Term
|   |   |   - number '4'
|   |   - ' '
|   +- operator '-'
|   +- ' '
|   - Factor
|       - Term
|           - number '5'
- ')'
Expression
+- Factor
|   - Term
|       - number '1'
+- ' '
+- operator '+'
+- ' '
- Factor
+- Term
|   - number '2'
+- ' '
+- operator '*'
+- ' '
- Term
+- '('
+- Expression
|   +- Factor
|   |   - Term
|   |       - number '3'
|   +- ' '
|   +- operator '+'
|   +- ' '
|   +- Factor
|   |   +- Term
|   |   |   - number '4'
|   |   - ' '
|   +- operator '-'
|   +- ' '
|   - Factor
|       - Term
|           - number '5'
- ')'
Expression
+- Factor
|   - Term
|       - number '1'
+- ' '
+- operator '+'
+- ' '
- Factor
+- Term
|   - number '2'
- ' '
Expression
+- Factor
|   - Term
|       - number '1'
+- ' '
+- operator '+'
+- ' '
- Factor
- Term
- number '2'
Expression
+- Factor
|   - Term
|       - number '1'
- ' '
Expression
- Factor
- Term
- number '1'`