Skip to content

ANTLR4 to TextMate grammar translation

Here is an example of what would happen if we translated the given ANTLR4 lexer grammar naively into a TextMate syntax-highlighting grammar:

ANTLR grammar actual expected DIV : '/' ; COMMENT : '//' ~[ \r\n ]* ; CMD : '/find' ; INT : [ 0 - 9 ]+ ; HEX : '0x' [ 0 - 9a-f ]+ ; ACCESS : ( 'read' | 'read write ' ) ; a / b // read me /find ijk 42 0x15af15 read x readwrite y a / b // read me /find ijk 42 0x15af15 read x readwrite y

We call this the shadowing problem. The core issue is that ANTLR-generated lexers differ in behavior from TextMate-based syntax highlighters when the input matches multiple rules:

  1. an ANTLR lexer picks the rule that matches the most characters from the input (→ longest-match)
  2. a TextMate tokenizer picks the first rule that finds any match (→ leftmost-match)

This also applies to alternatives (e.g. ('a' | 'ab')) within a single rule due to the way the Oniguruma regex engine works.

Just reordering the rules doesn’t help:

#\d{3} #\d+ #\d+ #\d{3} ? #0 #01 #012 #0123 #01234 #012345 #0 #01 #012 #0123 #01234 #012345 #0 #01 #012 #0123 #01234 #012345 × × ✔

// TODO: write about the solution