
A data structure used in computer science to represent the structure of a program or code snippet. It is a tree representation of the abstract syntactic structure of text, often source code, written in a formal language. Each node of the tree denotes a construct occurring in the text. It is sometimes called just a syntax tree. The syntax is "abstract" in the sense that it does not represent every detail appearing in the real syntax, but rather just the structural or content-related details.
How It Works
A parser turns code into an Abstract Syntax Tree (AST), where each node represents a construct like a function call, loop, or assignment. Instead of searching for a text string, an AST matcher searches for a shape: for example, any call to eval() with a variable argument. Because the match is structural, it finds every equivalent instance regardless of how the code is written, foo( x ), foo(x), and foo(\n x\n) all match the same pattern.
Applications
- Code Analysis: AST matching is precise and reliable, which is why it powers linters, codemods, refactoring tools, and static analysis.
- Pattern Detection: Tools like ast-grep, ESLint, and Semgrep use ASTs to detect patterns and rewrite code safely at scale.
- Example: To flag every console.log call, a regex might miss console .log() or match it inside a comment. An AST matcher targets the actual call expression node, catching all real calls and nothing else.
Advantages Over Text Search
- Precision: AST matching ignores formatting, whitespace, and variable names that don't affect meaning.
- Reliability: It finds every equivalent instance of a code construct, regardless of how the code is written.