mirror of
https://github.com/encounter/ghidra-cli.git
synced 2026-07-10 03:18:56 -07:00
This commit implements a complete Rust CLI tool for Ghidra reverse engineering, optimized for Claude Code and AI agents. Core Features: - Universal query command supporting all Ghidra data types (functions, strings, imports, exports, memory, etc.) - Advanced filter language with comparison, string, and logical operators - Multiple output formats (JSON, CSV, Table, minimal) optimized for LLM token efficiency - Field selection and pagination for precise data extraction - Windows-first design with cross-platform compatibility Architecture: - Filter parser using Pest grammar for robust expression parsing - Modular design with separate filter, format, query, and Ghidra integration layers - Headless Ghidra integration with built-in Python scripts for data extraction - Configuration system with environment variable and file support - Auto-detection of Ghidra installation on Windows LLM Optimizations: - Count-first workflow to check result sizes before fetching - Aggressive server-side filtering to reduce data transfer - Field selection to minimize token usage - Compact output formats (json-compact, minimal, ids) - Pagination support for large datasets Documentation: - Comprehensive README with examples and troubleshooting - Claude skill document (CLAUDE_SKILL.md) for agent integration - Subagent markdown (SUBAGENT.md) for Task tool integration - Inline code documentation and examples Commands Implemented: - ghidra query <data-type> - Universal query interface - ghidra import/analyze - Binary import and analysis - ghidra fn/strings/mem - Specialized command shortcuts - ghidra dump - Export data (imports, exports, functions, strings) - ghidra decompile - Function decompilation - ghidra project - Project management - ghidra config - Configuration management - ghidra init/doctor/version - Setup and diagnostics Built-in Ghidra Scripts: - Function listing with call graphs - Decompilation - String extraction - Import/Export tables - Memory map - Cross-references - Program information The CLI is designed to be succinct and efficient, with commands like: ghidra query functions --program=malware.exe --filter="size>1000 AND name~crypt" --format=json-compact Windows Support: - Auto-detection of Ghidra installation - Path handling for both Unix and Windows styles - Support for .exe, .dll, .sys formats This implementation provides a powerful, token-efficient interface for binary analysis that integrates seamlessly with Claude Code and other AI agents.
11 KiB
11 KiB
Ghidra Binary Analysis Subagent
Overview
This subagent specializes in reverse engineering binaries using Ghidra CLI. It provides efficient, token-optimized access to binary analysis capabilities for Claude Code and other AI agents.
When to Use This Subagent
Use this subagent when you need to:
- Analyze binary executables (PE, ELF, Mach-O)
- Reverse engineer malware or suspicious binaries
- Extract functions, strings, imports/exports from binaries
- Decompile functions to understand behavior
- Find specific patterns in binary code
- Analyze memory layout and structure
- Identify crypto functions, network operations, or file I/O
- Generate reports on binary capabilities
Capabilities
Data Extraction
- Functions: List, filter, and decompile functions
- Strings: Extract and search string literals
- Imports/Exports: Analyze external dependencies
- Memory Layout: View memory regions and permissions
- Symbols: Access symbol table
- Cross-References: Find call relationships
Analysis Features
- Universal Query System: Query any data type with powerful filters
- Advanced Filtering: Complex boolean expressions with field-level filtering
- Multiple Output Formats: JSON, CSV, Table, minimal (token-efficient)
- Decompilation: Convert assembly to C-like pseudocode
- Pattern Matching: Find specific code patterns and strings
LLM Optimizations
- Count-First Workflow: Check result sizes before fetching data
- Field Selection: Request only needed fields
- Aggressive Filtering: Pre-filter on Ghidra side
- Compact Formats: Minimal token usage with
json-compact - Pagination: Handle large datasets efficiently
Command Reference
Quick Start
# Import and analyze a binary
ghidra import <binary-path> --project=<project>
# Quick analysis (all-in-one)
ghidra quick <binary-path>
# Get program summary
ghidra summary --program=<binary>
Universal Query
# Query any data type
ghidra query <data-type> --program=<binary> [options]
# Data types: functions, strings, imports, exports, memory, symbols, xrefs
# Essential options:
--filter="<expression>" # Filter results
--fields=<list> # Select specific fields
--format=<format> # Output format (json, json-compact, table, count)
--limit=<n> # Max results
--count # Just return count
Common Queries
# List functions with filtering
ghidra query functions --program=<binary> \
--filter="size>1000 AND name~crypt" \
--fields=name,address,size \
--format=json-compact
# Find strings
ghidra query strings --program=<binary> \
--filter="value~http" \
--format=minimal
# List imports
ghidra dump imports --program=<binary> \
--filter="name~Crypt" \
--format=json-compact
# Get memory map
ghidra query memory --program=<binary> --format=table
Decompilation
# Decompile function by address
ghidra decompile 0x401000 --program=<binary>
# Decompile by name
ghidra decompile main --program=<binary>
# Compact output
ghidra fn decompile <addr> --program=<binary> --format=compact
Filter Language
Operators
Comparison: =, !=, >, >=, <, <=
String: ~ (contains), ^ (starts), $ (ends), =~ (regex)
Logical: AND, OR, NOT, ()
Special: EXISTS, IN [val1,val2]
Examples
# Exact match
name=malloc
# Numeric comparison
size>1000
# String matching (case-insensitive)
name~crypt
# Boolean logic
name~crypt AND size>500
(name~main OR name~start) AND NOT name^FUN_
# IN operator
name IN [malloc,free,realloc]
# Field existence
calls EXISTS
# Complex expression
size>=100 AND size<=1000 AND (name~crypt OR calls~Crypt)
Output Formats
count- Just the number (check result size)json-compact- Minimal JSON (best for LLMs)minimal- Addresses/names only (piping)ids- Just IDs (for further queries)table- Human-readable (display)json- Full JSON (complete data)
Best Practices
1. Count-First Pattern
Always check the result size before fetching data:
# Step 1: Count
ghidra query functions --program=<binary> --count
# Step 2: Refine filter if needed
ghidra query functions --program=<binary> \
--filter="NOT name^FUN_" \
--count
# Step 3: Fetch minimal data
ghidra query functions --program=<binary> \
--filter="NOT name^FUN_" \
--fields=name,address \
--format=json-compact \
--limit=50
2. Aggressive Filtering
Filter on Ghidra side, not in your code:
# GOOD: Pre-filter
ghidra query functions --program=<binary> \
--filter="size>1000 AND name~crypt"
# BAD: Fetch all, then filter
ghidra query functions --program=<binary> # Then filter in code
3. Field Selection
Request only what you need:
# Only name and address
ghidra query functions --program=<binary> \
--fields=name,address \
--format=json-compact
4. Use Compact Formats
Minimize token usage:
# For analysis: json-compact
--format=json-compact
# For display: table
--format=table
# For piping: minimal or ids
--format=ids
Analysis Workflows
Initial Reconnaissance
# 1. Get summary
ghidra summary --program=<binary>
# 2. Count functions
ghidra query functions --program=<binary> --count
# 3. Count named functions
ghidra query functions --program=<binary> \
--filter="NOT name^FUN_" --count
# 4. List key functions
ghidra query functions --program=<binary> \
--filter="NOT name^FUN_" \
--fields=name,address,size \
--format=json-compact \
--limit=20
Finding Suspicious Behavior
# Network operations
ghidra query imports --program=<binary> \
--filter="name~socket OR name~http OR name~inet" \
--format=json-compact
# File operations
ghidra query imports --program=<binary> \
--filter="name~File OR name~Read OR name~Write" \
--format=json-compact
# Process operations
ghidra query imports --program=<binary> \
--filter="name~Process OR name~Thread OR name~Exec" \
--format=json-compact
# Crypto operations
ghidra query imports --program=<binary> \
--filter="name~Crypt" \
--format=json-compact
String Analysis
# URLs
ghidra query strings --program=<binary> \
--filter="value~http" \
--format=json-compact
# Registry keys
ghidra query strings --program=<binary> \
--filter="value~HKEY OR value~Software" \
--format=json-compact
# Credentials
ghidra query strings --program=<binary> \
--filter="value~password OR value~username OR value~key" \
--format=json-compact
# Long strings (paths, URLs)
ghidra query strings --program=<binary> \
--filter="length>50" \
--format=json-compact
Deep Dive Analysis
# 1. Find interesting functions
ghidra query functions --program=<binary> \
--filter="name~crypt OR calls~Crypt" \
--fields=name,address \
--format=json-compact
# 2. Decompile each
ghidra decompile <address> --program=<binary>
# 3. Find cross-references
ghidra xref to <address> --program=<binary> \
--format=json-compact
# 4. Analyze callers
ghidra fn calls <address> --program=<binary> \
--format=json-compact
Common Patterns
Pattern: Find Entry Point
ghidra query functions --program=<binary> \
--filter="name~main OR name~WinMain OR name~DllMain" \
--format=json-compact
Pattern: Find Crypto Functions
# By name
ghidra query functions --program=<binary> \
--filter="name~crypt OR name~cipher OR name~aes OR name~rsa" \
--format=json-compact
# By imports
ghidra query imports --program=<binary> \
--filter="name~Crypt" \
--format=json-compact
Pattern: Find Large/Complex Functions
ghidra query functions --program=<binary> \
--filter="size>2000" \
--fields=name,address,size \
--sort=-size \
--limit=10 \
--format=json-compact
Pattern: Trace Function Calls
# What does function call?
ghidra fn calls <function> --program=<binary> \
--format=json-compact
# What calls this function?
ghidra fn xrefs <function> --program=<binary> \
--format=json-compact
Configuration
Set Defaults (avoid repeating --program)
# Set default program
ghidra set-default program <binary>
# Now you can omit --program
ghidra query functions --count
Environment Variables
# Windows
set GHIDRA_INSTALL_DIR=C:\ghidra\ghidra_11.0
set GHIDRA_DEFAULT_PROGRAM=malware.exe
# Unix
export GHIDRA_INSTALL_DIR=/opt/ghidra
export GHIDRA_DEFAULT_PROGRAM=malware.elf
Error Handling
Common Issues
- Ghidra not found: Run
ghidra initor setGHIDRA_INSTALL_DIR - Program not specified: Use
--program=<binary>or set default - Analysis timeout: Increase with
set GHIDRA_TIMEOUT=600 - Large result set: Use
--countfirst, then filter more aggressively
Troubleshooting
# Check installation
ghidra doctor
# Show configuration
ghidra config list
# List projects
ghidra project list
Performance Considerations
- Always count first - Prevents context overflow
- Filter aggressively - Reduce data before transfer
- Select minimal fields - Less data = fewer tokens
- Use compact formats -
json-compactis most efficient - Paginate results - Use
--limitfor large datasets - Cache results - Store commonly-used data in variables
Example: Complete Analysis
# Import binary
ghidra import suspicious.exe --project=analysis
# Set as default
ghidra set-default program suspicious.exe
# Overview
ghidra summary
# Count functions
ghidra query functions --count
# → 1247
# Named functions only
ghidra query functions --filter="NOT name^FUN_" --count
# → 89
# Get named functions
ghidra query functions \
--filter="NOT name^FUN_" \
--fields=name,address,size \
--format=json-compact
# Find suspicious imports
ghidra dump imports \
--filter="name~Exec OR name~Process OR name~Write" \
--format=json-compact
# Find URLs/IPs
ghidra query strings \
--filter="value~http OR value=~\"[0-9]{1,3}\\.[0-9]{1,3}\"" \
--format=json-compact
# Decompile interesting functions
ghidra decompile 0x401000
# Find what calls it
ghidra xref to 0x401000 --format=json-compact
Integration Tips
This subagent works best when:
- You have a binary file that needs analysis
- You need to understand malware behavior
- You're investigating suspicious executables
- You need to extract specific information (strings, functions, imports)
- You want to generate a report on binary capabilities
The subagent is optimized for:
- Token efficiency (minimal output)
- Fast queries (count-first pattern)
- Precise filtering (server-side pre-filtering)
- Flexible output (multiple formats)
- Automation-friendly (scriptable)
Limitations
- Requires Ghidra to be installed
- Windows path handling is primary (but cross-platform)
- Initial analysis can be slow for large binaries
- Decompilation quality depends on Ghidra's capabilities
- Complex analysis may require multiple queries
Support
- Run
ghidra doctorto check installation - See
README.mdfor full documentation - Check
CLAUDE_SKILL.mdfor detailed examples