Author SHA1 Message Date
Sylvestre Ledru 96762f26ca Add Default impl for GlobSet to satisfy clippy 2026-05-31 10:18:32 +02:00
Sylvestre Ledru c8dfef6563 Add pre-commit configuration 2026-05-31 10:11:57 +02:00
Sylvestre LedruandGitHub ddac723054 Merge pull request #13 from uutils/codspeed-wizard-1780213675759
Add CodSpeed performance benchmarks
2026-05-31 10:11:23 +02:00
codspeed-hq[bot]andGitHub 079619ee44 Add CodSpeed performance benchmarks 2026-05-31 07:59:22 +00:00
Sylvestre LedruandGitHub 2a6a3aba7e Add note about performance improvements needed 2026-05-30 19:25:27 +02:00
Sylvestre LedruandGitHub 98e6bb6f53 Merge pull request #10 from uutils/improv-cov
Improve the code coverage
2026-05-30 14:25:50 +02:00
Sylvestre Ledru da0ada8a37 test: use platform-specific expected message for nonexistent-file error 2026-05-30 11:33:47 +02:00
Sylvestre Ledru fcf46d8f56 test: cover -D skip on an explicit special-file argument 2026-05-30 11:28:04 +02:00
Sylvestre Ledru 55cb643545 test: cover strip_dot_prefix for implicit-cwd recursive search 2026-05-30 11:28:04 +02:00
Sylvestre Ledru 89aec4a45a test: cover -T line-number width on the file (not stdin) path 2026-05-30 11:28:03 +02:00
Sylvestre Ledru 6c5b7dc9e3 test: cover -L listing when the pattern matches one file 2026-05-30 11:28:03 +02:00
Sylvestre LedruandGitHub ff918e4a63 grep: strip trailing (os error N) from file error messages to match GNU (#6) 2026-05-29 22:24:20 +02:00
Sylvestre LedruandGitHub c50c0458cb fix: accept repeated options like GNU grep (#2)
GNU grep tolerates options given more than once (boolean flags are
idempotent, value options take the last occurrence), but clap rejected
them with "cannot be used multiple times" (exit 2). Set
args_override_self(true) so clap replaces rather than errors; args with
ArgAction::Append (-e/-f/--include/--exclude) still accumulate.

Found by the differential fuzzer (e.g. `grep -e .* -E -n -o -n -x`).

Add a regression test covering repeated booleans, repeated value options
(last wins), and that -e still accumulates patterns.
2026-05-29 22:21:27 +02:00
Sylvestre LedruandGitHub f7813500c9 Merge pull request #8 from uutils/gnu-test
run the gnu testsuite in the ci
2026-05-29 21:43:41 +02:00
Leonard Hecker 05d167b8d5 Fix support for Python-style named backreferences 2026-05-29 21:12:19 +02:00
Sylvestre LedruandGitHub 41b92b3cc1 Merge pull request #5 from uutils/fix-coverage-profraw-path
ci: use absolute LLVM_PROFILE_FILE path to fix 0% coverage
2026-05-29 20:46:10 +02:00
Sylvestre Ledru e885f66523 ci: post GNU grep testsuite comparison as a PR comment
Add a GnuComment workflow that runs after GnuTests completes on a pull
request, downloads the 'comment' artifact (PR number + comparison text), and
posts it as a PR comment. Mirrors ../sed's GnuComment workflow.
2026-05-29 18:43:11 +02:00
Sylvestre Ledru 10bfa405dc ci: add GnuTests workflow running the GNU grep testsuite
Add a GnuTests workflow that, on push and PR, fetches the GNU grep release
tarball, builds the Rust grep binary, runs util/run-gnu-testsuite.sh, and
uploads the JSON results. An aggregate job compares the run against the
reference summary from the default branch (util/compare_test_results.py,
borrowed from ../sed) and fails only on new, non-intermittent regressions;
known-flaky tests are listed in .github/workflows/ignore-intermittent.txt.
2026-05-29 18:43:11 +02:00
Sylvestre Ledru f0ae449d11 tests: run the GNU grep testsuite against uu_grep
Add util/fetch-gnu.sh (downloads the GNU grep 3.12 release tarball from
ftp.gnu.org) and util/run-gnu-testsuite.sh, which reuses the gnulib test
framework shipped in the tarball (tests/init.sh + init.cfg) and injects the
Rust grep binary via PATH, replicating tests/Makefile.am's TESTS_ENVIRONMENT.
Each test is classified by its gnulib exit code (0=PASS, 77=SKIP, else FAIL)
and results are emitted as JSON.

Modelled on ../sed (lightweight PATH-injection runner) and ../coreutils
(release-tarball fetch). Current baseline: 61 pass / 39 fail / 28 skip of 128
tests -- the failures quantify the remaining GNU-compatibility gap.
2026-05-29 18:43:11 +02:00
Sylvestre Ledru 79db36edbe add license headers 2026-05-29 10:12:38 +02:00
20 changed files with 2000 additions and 71 deletions
+83
View File
@@ -0,0 +1,83 @@
name: GnuComment
# Post the GNU grep testsuite comparison (produced by the GnuTests workflow)
# as a comment on the pull request.
on:
workflow_run:
workflows: ["GnuTests"]
types:
- completed
permissions: {}
jobs:
post-comment:
permissions:
actions: read # to list workflow run artifacts
pull-requests: write # to comment on the pr
runs-on: ubuntu-latest
if: >
github.event.workflow_run.event == 'pull_request'
steps:
- name: 'Download artifact'
uses: actions/github-script@v7
with:
script: |
// List all artifacts from GnuTests
var artifacts = await github.rest.actions.listWorkflowRunArtifacts({
owner: context.repo.owner,
repo: context.repo.repo,
run_id: ${{ github.event.workflow_run.id }},
});
// Download the "comment" artifact, which contains a PR number (NR) and result.txt
var matchArtifact = artifacts.data.artifacts.filter((artifact) => {
return artifact.name == "comment"
})[0];
if (!matchArtifact) {
console.log('No comment artifact found');
return;
}
var download = await github.rest.actions.downloadArtifact({
owner: context.repo.owner,
repo: context.repo.repo,
artifact_id: matchArtifact.id,
archive_format: 'zip',
});
var fs = require('fs');
fs.writeFileSync('${{ github.workspace }}/comment.zip', Buffer.from(download.data));
- run: unzip comment.zip || echo "Failed to unzip comment artifact"
- name: 'Comment on PR'
uses: actions/github-script@v7
with:
github-token: ${{ secrets.GITHUB_TOKEN }}
script: |
var fs = require('fs');
// Check if files exist
if (!fs.existsSync('./NR')) {
console.log('No NR file found, skipping comment');
return;
}
if (!fs.existsSync('./result.txt')) {
console.log('No result.txt file found, skipping comment');
return;
}
var issue_number = Number(fs.readFileSync('./NR'));
var content = fs.readFileSync('./result.txt');
if (content.toString().trim().length > 7) { // 7 because we have backquote + \n
await github.rest.issues.createComment({
owner: context.repo.owner,
repo: context.repo.repo,
issue_number: issue_number,
body: 'GNU grep testsuite comparison:\n```\n' + content + '```'
});
} else {
console.log('Comment content too short, skipping');
}
+180
View File
@@ -0,0 +1,180 @@
name: GnuTests
# Run the upstream GNU grep testsuite against the Rust grep implementation to
# track and guard byte-for-byte compatibility. See util/run-gnu-testsuite.sh.
on:
pull_request:
push:
branches:
- '*'
# End the current execution if there is a new changeset in the PR.
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: ${{ github.ref != 'refs/heads/main' }}
env:
DEFAULT_BRANCH: ${{ github.event.repository.default_branch }}
TEST_FULL_SUMMARY_FILE: 'grep-gnu-full-result.json'
jobs:
native:
name: Run GNU grep testsuite
runs-on: ubuntu-24.04
steps:
- name: Checkout code (grep)
uses: actions/checkout@v4
with:
path: 'grep'
persist-credentials: false
- uses: dtolnay/rust-toolchain@master
with:
toolchain: stable
- uses: Swatinem/rust-cache@v2
with:
workspaces: "./grep -> target"
- name: Fetch GNU grep testsuite
shell: bash
run: |
## Download and extract the upstream GNU grep release tarball
mkdir -p gnu.grep
cd gnu.grep
bash ../grep/util/fetch-gnu.sh
- name: Build Rust grep binary
shell: bash
run: |
cd 'grep'
cargo build --release
- name: Run GNU grep testsuite
shell: bash
run: |
cd 'grep'
export GNU_GREP_DIR="../gnu.grep"
./util/run-gnu-testsuite.sh --json-output "${{ env.TEST_FULL_SUMMARY_FILE }}" || true
- name: Upload full json results
uses: actions/upload-artifact@v4
with:
name: grep-gnu-full-result
path: grep/${{ env.TEST_FULL_SUMMARY_FILE }}
if-no-files-found: warn
aggregate:
needs: [native]
permissions:
actions: read
contents: read
pull-requests: read
name: Aggregate GNU test results
runs-on: ubuntu-24.04
steps:
- name: Initialize workflow variables
id: vars
shell: bash
run: |
## VARs setup
outputs() { step_id="${{ github.action }}"; for var in "$@" ; do echo steps.${step_id}.outputs.${var}="${!var}"; echo "${var}=${!var}" >> $GITHUB_OUTPUT; done; }
TEST_SUMMARY_FILE='grep-gnu-result.json'
outputs TEST_SUMMARY_FILE
- name: Checkout code (grep)
uses: actions/checkout@v4
with:
path: 'grep'
persist-credentials: false
- name: Retrieve reference artifacts
uses: dawidd6/action-download-artifact@v6
continue-on-error: true
with:
workflow: GnuTests.yml
branch: "${{ env.DEFAULT_BRANCH }}"
workflow_conclusion: completed
path: "reference"
if_no_artifact_found: warn
- name: Download full json results
uses: actions/download-artifact@v4
with:
name: grep-gnu-full-result
path: results
- name: Extract/summarize testing info
id: summary
shell: bash
run: |
## Extract/summarize testing info
outputs() { step_id="${{ github.action }}"; for var in "$@" ; do echo steps.${step_id}.outputs.${var}="${!var}"; echo "${var}=${!var}" >> $GITHUB_OUTPUT; done; }
RESULT_FILE="results/${{ env.TEST_FULL_SUMMARY_FILE }}"
if [[ ! -f "$RESULT_FILE" ]]; then
echo "::error ::Result file $RESULT_FILE not found"
find results -type f || true
exit 1
fi
TOTAL=$(jq -r '.summary.total // 0' "$RESULT_FILE")
PASS=$(jq -r '.summary.passed // 0' "$RESULT_FILE")
FAIL=$(jq -r '.summary.failed // 0' "$RESULT_FILE")
SKIP=$(jq -r '.summary.skipped // 0' "$RESULT_FILE")
output="GNU grep tests summary = TOTAL: $TOTAL / PASS: $PASS / FAIL: $FAIL / SKIP: $SKIP"
echo "${output}"
if [[ "$FAIL" -gt 0 ]]; then
echo "::warning ::${output}"
fi
outputs TOTAL PASS FAIL SKIP
- name: Compare test failures VS reference
shell: bash
run: |
## Compare current results against the reference summary from the default branch
REF_SUMMARY_FILE='reference/grep-gnu-full-result/${{ env.TEST_FULL_SUMMARY_FILE }}'
CURRENT_SUMMARY_FILE="results/${{ env.TEST_FULL_SUMMARY_FILE }}"
IGNORE_INTERMITTENT="grep/.github/workflows/ignore-intermittent.txt"
# Set up comment directory for the GnuComment workflow.
COMMENT_DIR="reference/comment"
mkdir -p ${COMMENT_DIR}
echo ${{ github.event.number }} > ${COMMENT_DIR}/NR
COMMENT_LOG="${COMMENT_DIR}/result.txt"
: > "${COMMENT_LOG}"
COMPARISON_RESULT=0
if test -f "${REF_SUMMARY_FILE}"; then
python3 grep/util/compare_test_results.py \
--ignore-file "${IGNORE_INTERMITTENT}" \
--output "${COMMENT_LOG}" \
"${CURRENT_SUMMARY_FILE}" "${REF_SUMMARY_FILE}" || COMPARISON_RESULT=$?
else
echo "::warning ::Skipping test comparison; no prior reference summary at '${REF_SUMMARY_FILE}'."
fi
if [ ${COMPARISON_RESULT} -eq 1 ]; then
echo "::error ::Found new non-intermittent test failures"
UPLOAD_EXIT=1
else
echo "::notice ::No new test failures detected"
UPLOAD_EXIT=0
fi
echo "UPLOAD_EXIT=${UPLOAD_EXIT}" >> $GITHUB_ENV
- name: Upload comparison log (for GnuComment workflow)
if: success() || failure()
uses: actions/upload-artifact@v4
with:
name: comment
path: reference/comment/
- name: Report test results
if: success() || failure()
shell: bash
run: |
echo "::notice ::GNU grep testsuite: TOTAL ${{ steps.summary.outputs.TOTAL }} / PASS ${{ steps.summary.outputs.PASS }} / FAIL ${{ steps.summary.outputs.FAIL }} / SKIP ${{ steps.summary.outputs.SKIP }}"
# Fail the job if the comparison found new non-intermittent regressions.
exit "${UPLOAD_EXIT:-0}"
+37
View File
@@ -0,0 +1,37 @@
name: CodSpeed
on:
push:
branches:
- "main"
pull_request:
# `workflow_dispatch` allows CodSpeed to trigger backtest
# performance analysis in order to generate initial data.
workflow_dispatch:
permissions:
contents: read
id-token: write
jobs:
codspeed:
name: Run benchmarks
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Setup Rust toolchain, cache and cargo-codspeed binary
uses: moonrepo/setup-rust@v0
with:
channel: stable
cache-target: release
bins: cargo-codspeed
- name: Build the benchmark target(s)
run: cargo codspeed build
- name: Run the benchmarks
uses: CodSpeedHQ/action@v4
with:
mode: simulation
run: cargo codspeed run
@@ -0,0 +1,7 @@
# List of intermittent test names to ignore in result comparisons
# Format: one test name per line, lines starting with # are comments
#
# Add test names that are known to be flaky or environment-dependent
# Example:
# basic_substitution
# line_address_test
+55
View File
@@ -0,0 +1,55 @@
# See https://pre-commit.com for more information
# See https://pre-commit.com/hooks.html for more hooks
exclude: ^tests/fixtures/
repos:
- repo: https://github.com/pre-commit/pre-commit-hooks
rev: v6.0.0
hooks:
- id: check-added-large-files
- id: check-executables-have-shebangs
- id: check-json
exclude: '\.vscode/(cSpell|extensions)\.json' # cSpell.json and extensions.json use comments
- id: check-shebang-scripts-are-executable
exclude: '.+\.rs' # would be triggered by #![some_attribute]
- id: check-symlinks
- id: check-toml
- id: check-yaml
args: [ --allow-multiple-documents ]
- id: destroyed-symlinks
- id: end-of-file-fixer
- id: mixed-line-ending
args: [ --fix=lf ]
- id: trailing-whitespace
- repo: local
hooks:
- id: rust-linting
name: Rust linting
description: Run cargo fmt on files included in the commit.
entry: cargo +stable fmt --
pass_filenames: true
types: [file, rust]
language: system
- id: rust-clippy
name: Rust clippy
description: Run cargo clippy on files included in the commit.
entry: cargo +stable clippy --workspace --all-targets --all-features -- -D warnings
pass_filenames: false
types: [file, rust]
language: system
- id: cargo-lock-check
name: Cargo.lock sync check
description: Ensure Cargo.lock and fuzz/Cargo.lock are up-to-date.
entry: bash -c 'for dir in . fuzz; do if [ -d "$dir" ]; then ( cd "$dir" && cargo fetch --quiet ); fi; done'
pass_filenames: false
files: 'Cargo\.(toml|lock)$'
language: system
- id: cspell
name: Code spell checker (cspell)
description: Run cspell to check for spelling errors (if available).
entry: bash -c 'if command -v cspell >/dev/null 2>&1; then cspell --no-must-find-files -- "$@"; else echo "cspell not found, skipping spell check"; exit 0; fi' --
pass_filenames: true
language: system
ci:
skip: [rust-linting, rust-clippy, cargo-lock-check, cspell]
Generated
+532 -11
View File
File diff suppressed because it is too large Load Diff
+5
View File
@@ -27,5 +27,10 @@ onig_sys = { version = "*", default-features = false }
uucore = "0.8.0"
walkdir = "2.5"
[[bench]]
name = "grep_bench"
harness = false
[dev-dependencies]
criterion = { version = "4.7.0", package = "codspeed-criterion-compat" }
uutests = "0.8.0"
+6
View File
@@ -4,6 +4,7 @@
[![dependency status](https://deps.rs/repo/github/uutils/grep/status.svg)](https://deps.rs/repo/github/uutils/grep)
[![CodeCov](https://codecov.io/gh/uutils/grep/branch/main/graph/badge.svg)](https://codecov.io/gh/uutils/grep)
[![CodSpeed](https://img.shields.io/endpoint?url=https://codspeed.io/badge.json)](https://codspeed.io/uutils/grep?utm_source=badge)
# Grep, now in Rust
@@ -29,10 +30,15 @@ cargo build --release
cargo test
```
## Pre-commit hooks
This project uses [pre-commit](https://pre-commit.com); run `pre-commit install` to enable the git hooks.
## Known Issues
* Does not take `LANG`, etc., into account for handling file encodings (non-UTF8 matches are treated as binary)
* No localization support yet
* Performances need to be improved
## Contributing
+288
View File
@@ -0,0 +1,288 @@
// This file is part of the uutils grep package.
//
// For the full copyright and license information, please view the LICENSE
// file that was distributed with this source code.
use criterion::{Criterion, black_box, criterion_group, criterion_main};
use uu_grep::matcher::Matcher;
use uu_grep::{BinaryMode, ColorConfig, Config, DeviceMode, DirectoryMode, GlobSet, RegexMode};
fn make_config<'a>(
patterns: &'a [&'a str],
regex_mode: RegexMode,
ignore_case: bool,
invert_match: bool,
word_regexp: bool,
) -> Config<'a> {
Config {
directory_mode: DirectoryMode::Read,
device_mode: DeviceMode::Default,
follow_symlinks: false,
include_globs: GlobSet::new(),
exclude_globs: GlobSet::new(),
exclude_dir_globs: GlobSet::new(),
label: "(standard input)",
#[cfg(windows)]
strip_cr: false,
binary_mode: BinaryMode::Binary,
max_count: None,
before_context: 0,
after_context: 0,
has_context: false,
patterns,
regex_mode,
ignore_case,
invert_match,
word_regexp,
line_regexp: false,
quiet: true,
count: false,
show_filename: false,
files_with_matches: false,
files_without_match: false,
only_matching: false,
byte_offset: false,
line_number: false,
initial_tab: false,
null_separator: false,
null_data: false,
line_buffered: false,
no_messages: true,
group_separator: None,
use_color: false,
color_config: ColorConfig {
matched_selected: "",
matched_context: "",
filename: "",
line_number: "",
byte_offset: "",
separator: "",
selected_line: "",
context_line: "",
reverse_video: false,
no_erase: false,
},
}
}
fn bench_compile(c: &mut Criterion) {
let mut group = c.benchmark_group("compile");
group.bench_function("fixed_string", |b| {
b.iter(|| {
let patterns: &[&str] = &["hello world"];
let config = make_config(patterns, RegexMode::Fixed, false, false, false);
let matcher = Matcher::compile(black_box(&config)).unwrap();
let _ = black_box(&matcher);
})
});
group.bench_function("basic_regex", |b| {
b.iter(|| {
let patterns: &[&str] = &[r"[0-9]\{4\}-[0-9]\{2\}-[0-9]\{2\}"];
let config = make_config(patterns, RegexMode::Basic, false, false, false);
let matcher = Matcher::compile(black_box(&config)).unwrap();
let _ = black_box(&matcher);
})
});
group.bench_function("extended_regex", |b| {
b.iter(|| {
let patterns: &[&str] = &[r"[0-9]{4}-[0-9]{2}-[0-9]{2}"];
let config = make_config(patterns, RegexMode::Extended, false, false, false);
let matcher = Matcher::compile(black_box(&config)).unwrap();
let _ = black_box(&matcher);
})
});
group.bench_function("perl_regex", |b| {
b.iter(|| {
let patterns: &[&str] = &[r"\d{4}-\d{2}-\d{2}\s+\d{2}:\d{2}:\d{2}"];
let config = make_config(patterns, RegexMode::Perl, false, false, false);
let matcher = Matcher::compile(black_box(&config)).unwrap();
let _ = black_box(&matcher);
})
});
group.bench_function("multiple_patterns", |b| {
b.iter(|| {
let patterns: &[&str] = &["error", "warning", "critical", "fatal", "panic"];
let config = make_config(patterns, RegexMode::Fixed, false, false, false);
let matcher = Matcher::compile(black_box(&config)).unwrap();
let _ = black_box(&matcher);
})
});
group.finish();
}
fn bench_match(c: &mut Criterion) {
let mut group = c.benchmark_group("match");
// Fixed string match - hit
{
let patterns: &[&str] = &["ERROR"];
let config = make_config(patterns, RegexMode::Fixed, false, false, false);
let matcher = Matcher::compile(&config).unwrap();
let line = b"2024-01-15 10:30:45 ERROR: Connection timeout on server-42";
group.bench_function("fixed_string_hit", |b| {
b.iter(|| black_box(matcher.match_line(black_box(line))))
});
}
// Fixed string match - miss
{
let patterns: &[&str] = &["CRITICAL"];
let config = make_config(patterns, RegexMode::Fixed, false, false, false);
let matcher = Matcher::compile(&config).unwrap();
let line = b"2024-01-15 10:30:45 INFO: Server started successfully";
group.bench_function("fixed_string_miss", |b| {
b.iter(|| black_box(matcher.match_line(black_box(line))))
});
}
// Extended regex match
{
let patterns: &[&str] = &[r"[0-9]{4}-[0-9]{2}-[0-9]{2}"];
let config = make_config(patterns, RegexMode::Extended, false, false, false);
let matcher = Matcher::compile(&config).unwrap();
let line = b"2024-01-15 10:30:45 ERROR: Connection timeout";
group.bench_function("extended_regex_hit", |b| {
b.iter(|| black_box(matcher.match_line(black_box(line))))
});
}
// Case-insensitive match
{
let patterns: &[&str] = &["error"];
let config = make_config(patterns, RegexMode::Fixed, true, false, false);
let matcher = Matcher::compile(&config).unwrap();
let line = b"2024-01-15 10:30:45 ERROR: Connection timeout";
group.bench_function("case_insensitive_hit", |b| {
b.iter(|| black_box(matcher.match_line(black_box(line))))
});
}
// Inverted match
{
let patterns: &[&str] = &["ERROR"];
let config = make_config(patterns, RegexMode::Fixed, false, true, false);
let matcher = Matcher::compile(&config).unwrap();
let line = b"2024-01-15 10:30:45 INFO: Server started successfully";
group.bench_function("inverted_match", |b| {
b.iter(|| black_box(matcher.match_line(black_box(line))))
});
}
// Word boundary match
{
let patterns: &[&str] = &["error"];
let config = make_config(patterns, RegexMode::Fixed, true, false, true);
let matcher = Matcher::compile(&config).unwrap();
let line = b"2024-01-15 10:30:45 error: Connection timeout";
group.bench_function("word_boundary_hit", |b| {
b.iter(|| black_box(matcher.match_line(black_box(line))))
});
}
// Multiple patterns
{
let patterns: &[&str] = &["error", "warning", "critical", "fatal", "panic"];
let config = make_config(patterns, RegexMode::Fixed, true, false, false);
let matcher = Matcher::compile(&config).unwrap();
let line = b"2024-01-15 10:30:45 WARNING: High memory usage detected on node-7";
group.bench_function("multi_pattern_hit", |b| {
b.iter(|| black_box(matcher.match_line(black_box(line))))
});
}
// Long line
{
let patterns: &[&str] = &["needle"];
let config = make_config(patterns, RegexMode::Fixed, false, false, false);
let matcher = Matcher::compile(&config).unwrap();
let mut long_line = "a".repeat(5000);
long_line.push_str("needle");
long_line.push_str(&"b".repeat(5000));
let long_line_bytes = long_line.into_bytes();
group.bench_function("long_line_hit", |b| {
b.iter(|| black_box(matcher.match_line(black_box(&long_line_bytes))))
});
}
group.finish();
}
fn bench_throughput(c: &mut Criterion) {
let mut group = c.benchmark_group("throughput");
// Simulate processing many lines (like searching a log file)
let lines: Vec<Vec<u8>> = (0..1000)
.map(|i| {
if i % 50 == 0 {
format!(
"2024-01-15 10:30:{:02} ERROR: Connection timeout on server-{}",
i % 60,
i
)
.into_bytes()
} else {
format!(
"2024-01-15 10:30:{:02} INFO: Request processed in {}ms",
i % 60,
i * 3
)
.into_bytes()
}
})
.collect();
{
let patterns: &[&str] = &["ERROR"];
let config = make_config(patterns, RegexMode::Fixed, false, false, false);
let matcher = Matcher::compile(&config).unwrap();
group.bench_function("scan_1000_lines_fixed", |b| {
b.iter(|| {
let mut matches = 0u64;
for line in &lines {
if matcher.match_line(black_box(line)).is_some() {
matches += 1;
}
}
black_box(matches)
})
});
}
{
let patterns: &[&str] = &[r"[0-9]+ *ms"];
let config = make_config(patterns, RegexMode::Extended, false, false, false);
let matcher = Matcher::compile(&config).unwrap();
group.bench_function("scan_1000_lines_regex", |b| {
b.iter(|| {
let mut matches = 0u64;
for line in &lines {
if matcher.match_line(black_box(line)).is_some() {
matches += 1;
}
}
black_box(matches)
})
});
}
group.finish();
}
criterion_group!(benches, bench_compile, bench_match, bench_throughput);
criterion_main!(benches);
+5
View File
@@ -1,3 +1,8 @@
// This file is part of the uutils grep package.
//
// For the full copyright and license information, please view the LICENSE
// file that was distributed with this source code.
pub struct LineView<'a> {
/// Line content (without the terminator).
pub line: &'a [u8],
+88 -56
View File
@@ -1,6 +1,14 @@
mod context_buffer;
mod line_buffer;
mod matcher;
// This file is part of the uutils grep package.
//
// For the full copyright and license information, please view the LICENSE
// file that was distributed with this source code.
#[doc(hidden)]
pub mod context_buffer;
#[doc(hidden)]
pub mod line_buffer;
#[doc(hidden)]
pub mod matcher;
mod output;
mod searcher;
@@ -15,7 +23,8 @@ use std::path::Path;
use uucore::error::{FromIo, UResult, USimpleError};
#[derive(Clone, Copy, PartialEq, Eq)]
enum RegexMode {
#[doc(hidden)]
pub enum RegexMode {
Fixed,
Basic,
Extended,
@@ -23,7 +32,8 @@ enum RegexMode {
}
#[derive(Clone, Copy, PartialEq, Eq)]
enum BinaryMode {
#[doc(hidden)]
pub enum BinaryMode {
Binary,
Text,
WithoutMatch,
@@ -37,79 +47,84 @@ enum ColorMode {
}
#[derive(Clone, Copy, PartialEq, Eq)]
enum DirectoryMode {
#[doc(hidden)]
pub enum DirectoryMode {
Read,
Skip,
Recurse,
}
#[derive(Clone, Copy, PartialEq, Eq)]
enum DeviceMode {
#[doc(hidden)]
pub enum DeviceMode {
Default,
Read,
Skip,
}
struct ColorConfig<'a> {
matched_selected: &'a str,
matched_context: &'a str,
filename: &'a str,
line_number: &'a str,
byte_offset: &'a str,
separator: &'a str,
selected_line: &'a str,
context_line: &'a str,
#[doc(hidden)]
pub struct ColorConfig<'a> {
pub matched_selected: &'a str,
pub matched_context: &'a str,
pub filename: &'a str,
pub line_number: &'a str,
pub byte_offset: &'a str,
pub separator: &'a str,
pub selected_line: &'a str,
pub context_line: &'a str,
reverse_video: bool,
no_erase: bool,
pub reverse_video: bool,
pub no_erase: bool,
}
struct GlobSet {
#[doc(hidden)]
pub struct GlobSet {
patterns: Vec<glob::Pattern>,
}
struct Config<'a> {
#[doc(hidden)]
pub struct Config<'a> {
// Searcher
directory_mode: DirectoryMode,
device_mode: DeviceMode,
follow_symlinks: bool,
include_globs: GlobSet,
exclude_globs: GlobSet,
exclude_dir_globs: GlobSet,
label: &'a str,
pub directory_mode: DirectoryMode,
pub device_mode: DeviceMode,
pub follow_symlinks: bool,
pub include_globs: GlobSet,
pub exclude_globs: GlobSet,
pub exclude_dir_globs: GlobSet,
pub label: &'a str,
#[cfg(windows)]
strip_cr: bool,
binary_mode: BinaryMode,
max_count: Option<u64>,
before_context: usize,
after_context: usize,
has_context: bool,
pub strip_cr: bool,
pub binary_mode: BinaryMode,
pub max_count: Option<u64>,
pub before_context: usize,
pub after_context: usize,
pub has_context: bool,
// Matcher
patterns: &'a [&'a str],
regex_mode: RegexMode,
ignore_case: bool,
invert_match: bool,
word_regexp: bool,
line_regexp: bool,
pub patterns: &'a [&'a str],
pub regex_mode: RegexMode,
pub ignore_case: bool,
pub invert_match: bool,
pub word_regexp: bool,
pub line_regexp: bool,
// Output
quiet: bool,
count: bool,
show_filename: bool,
files_with_matches: bool,
files_without_match: bool,
only_matching: bool,
byte_offset: bool,
line_number: bool,
initial_tab: bool,
null_separator: bool,
null_data: bool,
line_buffered: bool,
no_messages: bool,
group_separator: Option<&'a str>,
use_color: bool,
color_config: ColorConfig<'a>,
pub quiet: bool,
pub count: bool,
pub show_filename: bool,
pub files_with_matches: bool,
pub files_without_match: bool,
pub only_matching: bool,
pub byte_offset: bool,
pub line_number: bool,
pub initial_tab: bool,
pub null_separator: bool,
pub null_data: bool,
pub line_buffered: bool,
pub no_messages: bool,
pub group_separator: Option<&'a str>,
pub use_color: bool,
pub color_config: ColorConfig<'a>,
}
#[uucore::main(no_signals)]
@@ -416,6 +431,10 @@ pub fn uu_app() -> Command {
.about("Search for PATTERNS in each FILE.")
.disable_help_flag(true)
.disable_version_flag(true)
// GNU grep accepts repeated options (booleans are idempotent, value
// options take the last); make clap replace rather than error. Args
// with ArgAction::Append (e.g. -e/-f/--include) still accumulate.
.args_override_self(true)
.after_help(
"When FILE is '-', read standard input. If no FILE is given, read standard \
input, but with -r, recursively search the working directory instead. With \
@@ -845,8 +864,21 @@ fn expand_num_shorthand(args: impl Iterator<Item = OsString>) -> Vec<OsString> {
out
}
impl Default for GlobSet {
fn default() -> Self {
Self::new()
}
}
impl GlobSet {
fn with_capacity(capacity: usize) -> Self {
/// Create an empty GlobSet.
pub fn new() -> Self {
Self {
patterns: Vec::new(),
}
}
pub fn with_capacity(capacity: usize) -> Self {
Self {
patterns: Vec::with_capacity(capacity),
}
+5
View File
@@ -1,3 +1,8 @@
// This file is part of the uutils grep package.
//
// For the full copyright and license information, please view the LICENSE
// file that was distributed with this source code.
use memchr::memchr;
use std::fs::File;
use std::io::{self, Read as _};
+5
View File
@@ -1 +1,6 @@
// This file is part of the uutils grep package.
//
// For the full copyright and license information, please view the LICENSE
// file that was distributed with this source code.
uucore::bin!(uu_grep);
+23 -3
View File
@@ -1,5 +1,13 @@
// This file is part of the uutils grep package.
//
// For the full copyright and license information, please view the LICENSE
// file that was distributed with this source code.
use crate::{Config, RegexMode};
use onig::{EncodedBytes, Regex, RegexOptions, Region, SearchOptions, Syntax, SyntaxBehavior};
use onig::{
EncodedBytes, Regex, RegexOptions, Region, SearchOptions, Syntax, SyntaxBehavior,
SyntaxOperator,
};
use onig_sys::{OnigEncCtype_ONIGENC_CTYPE_WORD, OnigEncodingUTF8};
use uucore::error::{UResult, USimpleError};
@@ -201,11 +209,23 @@ impl CompiledPattern {
RegexMode::Fixed => Syntax::asis(),
RegexMode::Basic => Syntax::grep(),
RegexMode::Extended => Syntax::gnu_regex(),
RegexMode::Perl => Syntax::perl(),
RegexMode::Perl => Syntax::perl_ng(),
};
if !matches!(config.regex_mode, RegexMode::Fixed) {
if config.regex_mode != RegexMode::Fixed {
// GNU grep supports `{,n}` as an alias for `{0,n}`.
syntax.enable_behavior(SyntaxBehavior::SYNTAX_BEHAVIOR_ALLOW_INTERVAL_LOW_ABBREV);
}
if config.regex_mode == RegexMode::Perl {
// GNU grep supports `(?P<name>...)`.
// Unfortunately, the onig crate defines the OP2 flag without the
// necessary <<32 bit shift, so we need to hotpatch that here.
const _: () =
assert!(SyntaxOperator::SYNTAX_OPERATOR_QMARK_CAPITAL_P_NAME.bits() == 0x80000000);
const FIXED: SyntaxOperator = SyntaxOperator::from_bits_retain(
SyntaxOperator::SYNTAX_OPERATOR_QMARK_CAPITAL_P_NAME.bits() << 32,
);
syntax.enable_operators(FIXED);
}
let mut options = RegexOptions::REGEX_OPTION_NONE;
if config.ignore_case {
+12 -1
View File
@@ -1,8 +1,14 @@
// This file is part of the uutils grep package.
//
// For the full copyright and license information, please view the LICENSE
// file that was distributed with this source code.
use crate::Config;
use crate::context_buffer::LineView;
use std::ffi::OsStr;
use std::io::{self, BufWriter, StdoutLock, Write};
use std::path::Path;
use uucore::error::strip_errno;
#[cfg(target_pointer_width = "64")]
const BUF_SIZE: usize = 128 * 1024;
@@ -190,7 +196,12 @@ impl<'a> OutputWriter<'a> {
/// Write an IO error to stderr.
pub fn report_io_error(&self, label: &OsStr, err: &io::Error) {
if !self.config.no_messages && !self.config.quiet {
eprintln!("grep: {label}: {err}", label = label.to_string_lossy());
// Strip the trailing " (os error XX)" so the message matches GNU grep.
eprintln!(
"grep: {label}: {err}",
label = label.to_string_lossy(),
err = strip_errno(err)
);
}
}
+5
View File
@@ -1,3 +1,8 @@
// This file is part of the uutils grep package.
//
// For the full copyright and license information, please view the LICENSE
// file that was distributed with this source code.
use crate::context_buffer::{ContextBuffer, LineView};
use crate::line_buffer::LineBuffer;
use crate::matcher::Matcher;
+115
View File
@@ -440,6 +440,17 @@ fn files_with_and_without_matches() {
.fails_with_code(1)
.stdout_is("hit\nmiss\n");
// -L with a pattern that DOES match in one file: the matching file is
// excluded from the listing, so only the non-matching file is printed.
// This exercises the early-return in `session_handle_match` taken when
// `files_without_match` is set and a match is found (src/searcher.rs).
let (scene, mut c) = ucmd();
scene.fixtures.write("hit", "yes\n");
scene.fixtures.write("miss", "no\n");
c.args(&["-L", "yes", "hit", "miss"])
.succeeds()
.stdout_is("miss\n");
// -l early-exits after the first match. Verify it doesn't print twice.
// when the file has many.
let (scene, mut c) = ucmd();
@@ -615,6 +626,15 @@ fn line_number_and_byte_offset_prefixes() {
.pipe_in("x\n")
.succeeds()
.stdout_contains("1:\tx\n");
// -T against a real file (not stdin): the line-number field width is
// derived from the file size, which only happens on the `File` path in
// `process_file` (src/searcher.rs), not the stdin path.
let (scene, mut c) = ucmd();
scene.fixtures.write("f", "x\n");
c.args(&["-T", "-n", "x", "f"])
.succeeds()
.stdout_contains("1:\tx\n");
}
#[test]
@@ -883,6 +903,25 @@ fn recursive_no_file_defaults_to_cwd_not_stdin() {
.stdout_contains("a.txt");
}
#[test]
fn recursive_implicit_cwd_strips_dot_prefix() {
// `-r` with no path argument searches the implicit ".", so reported paths
// come back as "./only.txt" (or ".\\only.txt" on Windows). GNU strips that
// leading prefix; `strip_dot_prefix` in src/searcher.rs must do the same.
// A single file keeps the output deterministic and lets us assert the exact
// line, which `stdout_contains` in `recursive_no_file_defaults_to_cwd_not_stdin`
// cannot (it would also pass with a leaked "./" prefix).
let (scene, _) = ucmd();
scene.fixtures.mkdir_all("flat");
scene.fixtures.write("flat/only.txt", "grep me\n");
let mut c = scene.cmd(env!("CARGO_BIN_EXE_grep"));
c.current_dir(scene.fixtures.plus("flat"))
.args(&["-r", "grep"])
.succeeds()
.stdout_is("only.txt:grep me\n");
}
#[test]
fn recursive_with_include_exclude() {
let (scene, _) = ucmd();
@@ -1022,6 +1061,30 @@ fn recursive_skips_fifos_by_default() {
.stdout_does_not_contain("fifo");
}
#[cfg(unix)]
#[test]
fn device_skip_on_explicit_special_file_arg() {
use std::process::Command;
// A special file (FIFO) named *directly* as an argument, not via recursion.
// With `-D skip` it must be dropped without reading (reading would block
// forever, so the test returning at all proves it was skipped). This covers
// the top-level special-file branch in `process_path` and `is_special_file`
// (src/searcher.rs), distinct from the recursive FIFO path.
let (scene, _) = ucmd();
let fifo_path = scene.fixtures.plus("fifo");
let status = Command::new("mkfifo")
.arg(&fifo_path)
.status()
.expect("mkfifo failed");
assert!(status.success(), "could not create FIFO");
let mut c = scene.cmd(env!("CARGO_BIN_EXE_grep"));
c.args(&["-D", "skip", "grep", "fifo"])
.fails_with_code(1)
.no_output();
}
#[test]
fn nonexistent_file_is_error() {
let (_s, mut c) = ucmd();
@@ -1030,6 +1093,23 @@ fn nonexistent_file_is_error() {
.stderr_contains("does-not-exist");
}
#[test]
fn nonexistent_file_error_has_no_os_error_suffix() {
// GNU prints "grep: <file>: No such file or directory" with no
// " (os error 2)" suffix; strip_errno keeps us byte-compatible. The
// underlying OS message text differs on Windows, but in both cases the
// trailing " (os error N)" must be absent.
#[cfg(not(windows))]
let expected = "grep: does-not-exist: No such file or directory\n";
#[cfg(windows)]
let expected = "grep: does-not-exist: The system cannot find the file specified.\n";
let (_s, mut c) = ucmd();
c.args(&["x", "does-not-exist"])
.fails_with_code(2)
.stderr_is(expected);
}
#[test]
fn dash_argument_means_stdin() {
let (_s, mut c) = ucmd();
@@ -1138,3 +1218,38 @@ fn help_and_version() {
.succeeds()
.stdout_contains(env!("CARGO_PKG_VERSION"));
}
#[test]
fn repeated_options_are_accepted() {
// GNU grep tolerates options given more than once: boolean flags are
// idempotent and value options take the last occurrence. clap would
// otherwise error with "cannot be used multiple times".
// Repeated boolean flags are a no-op (not an error).
let (_s, mut c) = ucmd();
c.args(&["-n", "-n", "a"])
.pipe_in("abc\n")
.succeeds()
.stdout_only("1:abc\n");
// Mixed repeated booleans behave like a single occurrence.
let (_s, mut c) = ucmd();
c.args(&["-i", "-i", "abc"])
.pipe_in("ABC\n")
.succeeds()
.stdout_only("ABC\n");
// Repeated value options take the last value (here: -m 1 wins).
let (_s, mut c) = ucmd();
c.args(&["-m", "5", "-m", "1", "x"])
.pipe_in("x\nx\nx\n")
.succeeds()
.stdout_only("x\n");
// -e (ArgAction::Append) must still accumulate every pattern.
let (_s, mut c) = ucmd();
c.args(&["-e", "a", "-e", "b"])
.pipe_in("a\nb\nc\n")
.succeeds()
.stdout_only("a\nb\n");
}
+173
View File
@@ -0,0 +1,173 @@
#!/usr/bin/env python3
"""
Compare the current GNU test results to the last results gathered from the main branch to
highlight if a PR is making the results better/worse.
Don't exit with error code if all failing tests are in the ignore-intermittent.txt list.
"""
import json
import sys
import argparse
from pathlib import Path
def load_ignore_list(ignore_file):
"""Load list of intermittent test names to ignore from file."""
ignore_set = set()
if ignore_file and Path(ignore_file).exists():
with open(ignore_file, "r") as f:
for line in f:
line = line.strip()
if line and not line.startswith("#"):
ignore_set.add(line)
return ignore_set
def extract_test_results(json_data):
"""Extract test results from JSON data."""
if not json_data or "summary" not in json_data:
return {"total": 0, "passed": 0, "failed": 0, "skipped": 0}, []
summary = json_data["summary"]
tests = json_data.get("tests", [])
# Extract failed test names
failed_tests = []
for test in tests:
if test.get("status") == "FAIL":
failed_tests.append(test.get("name", "unknown"))
return summary, failed_tests
def compare_results(current_file, reference_file, ignore_file=None, output_file=None):
"""Compare current results with reference results."""
# Load ignore list
ignore_set = load_ignore_list(ignore_file)
# Load JSON files
try:
with open(current_file, "r") as f:
current_data = json.load(f)
current_summary, current_failed = extract_test_results(current_data)
except Exception as e:
print(f"Error loading current results: {e}")
return 1
try:
with open(reference_file, "r") as f:
reference_data = json.load(f)
reference_summary, reference_failed = extract_test_results(reference_data)
except Exception as e:
print(f"Error loading reference results: {e}")
return 1
# Calculate differences
pass_diff = int(current_summary.get("passed", 0)) - int(
reference_summary.get("passed", 0)
)
fail_diff = int(current_summary.get("failed", 0)) - int(
reference_summary.get("failed", 0)
)
total_diff = int(current_summary.get("total", 0)) - int(
reference_summary.get("total", 0)
)
# Find new failures and improvements
current_failed_set = set(current_failed)
reference_failed_set = set(reference_failed)
new_failures = current_failed_set - reference_failed_set
improvements = reference_failed_set - current_failed_set
# Filter out intermittent failures
non_intermittent_new_failures = new_failures - ignore_set
# Check if results are identical (no changes)
no_changes = (
pass_diff == 0
and fail_diff == 0
and total_diff == 0
and not new_failures
and not improvements
)
# If no changes, write empty output to prevent comment posting
if no_changes:
with open(output_file, "w") as f:
f.write("")
return 0
# Prepare output message
output_lines = []
# Show current vs reference numbers for debugging
output_lines.append("Test results comparison:")
output_lines.append(
f" Current: TOTAL: {current_summary.get('total', 0)} / PASSED: {current_summary.get('passed', 0)} / FAILED: {current_summary.get('failed', 0)} / SKIPPED: {current_summary.get('skipped', 0)}"
)
output_lines.append(
f" Reference: TOTAL: {reference_summary.get('total', 0)} / PASSED: {reference_summary.get('passed', 0)} / FAILED: {reference_summary.get('failed', 0)} / SKIPPED: {reference_summary.get('skipped', 0)}"
)
output_lines.append("")
# Summary of changes
if pass_diff != 0 or fail_diff != 0 or total_diff != 0:
output_lines.append("Changes from main branch:")
output_lines.append(f" TOTAL: {total_diff:+d}")
output_lines.append(f" PASSED: {pass_diff:+d}")
output_lines.append(f" FAILED: {fail_diff:+d}")
output_lines.append("")
# New failures
if new_failures:
output_lines.append(f"New test failures ({len(new_failures)}):")
for test in sorted(new_failures):
if test in ignore_set:
output_lines.append(f" - {test} (intermittent)")
else:
output_lines.append(f" - {test}")
output_lines.append("")
# Improvements
if improvements:
output_lines.append(f"Test improvements ({len(improvements)}):")
for test in sorted(improvements):
output_lines.append(f" + {test}")
output_lines.append("")
# Write output
output_text = "\n".join(output_lines)
if output_file:
with open(output_file, "w") as f:
f.write(output_text)
else:
print(output_text)
# Return appropriate exit code
if non_intermittent_new_failures:
print(
f"ERROR: Found {len(non_intermittent_new_failures)} new non-intermittent test failures"
)
return 1
return 0
def main():
parser = argparse.ArgumentParser(description="Compare GNU test results")
parser.add_argument("current", help="Current test results JSON file")
parser.add_argument("reference", help="Reference test results JSON file")
parser.add_argument(
"--ignore-file", help="File containing intermittent test names to ignore"
)
parser.add_argument("--output", help="Output file for comparison results")
args = parser.parse_args()
return compare_results(args.current, args.reference, args.ignore_file, args.output)
if __name__ == "__main__":
sys.exit(main())
+17
View File
@@ -0,0 +1,17 @@
#!/bin/bash -e
# This file is part of the uutils grep package.
#
# For the full copyright and license information, please view the LICENSE
# file that was distributed with this source code.
#
# Download and extract the upstream GNU grep release tarball into the current
# directory. Run it from an (empty) directory that will hold the GNU grep tree,
# e.g.:
#
# mkdir -p ../gnu.grep && (cd ../gnu.grep && bash ../grep/util/fetch-gnu.sh)
#
# The extracted tree ships a ready-to-use gnulib test framework under tests/
# (init.sh + init.cfg + the extensionless test scripts), which
# util/run-gnu-testsuite.sh drives against the Rust grep binary.
ver="3.12"
curl -L "https://ftp.gnu.org/gnu/grep/grep-${ver}.tar.xz" | tar --strip-components=1 -xJf -
+359
View File
@@ -0,0 +1,359 @@
#!/bin/bash
# This file is part of the uutils grep package.
#
# For the full copyright and license information, please view the LICENSE
# file that was distributed with this source code.
#
# Run the upstream GNU grep testsuite against the Rust grep implementation.
#
# Unlike GNU coreutils, we do *not* build GNU grep here. Instead we reuse the
# gnulib test framework (tests/init.sh + tests/init.cfg) shipped in the GNU grep
# release tarball and inject our Rust `grep` binary via PATH, replicating the
# environment that tests/Makefile.am's TESTS_ENVIRONMENT would normally set up.
# Each test is classified by its gnulib exit code: 0 = PASS, 77 = SKIP, anything
# else = FAIL (timeouts and framework failures count as FAIL).
#
# Get the GNU grep sources with:
# mkdir -p ../gnu.grep && (cd ../gnu.grep && bash ../grep/util/fetch-gnu.sh)
#
# Usage: ./util/run-gnu-testsuite.sh [options]
#
# Options:
# -h, --help Show this help message
# -v, --verbose Show diagnostics for failing/skipped tests
# -q, --quiet Only print failures and the final summary
# --json-output FILE Write results to FILE as JSON
#
# Environment variables:
# GNU_GREP_DIR Path to the extracted GNU grep source tree
# (default: ../gnu.grep)
# RUN_EXPENSIVE_TESTS Set to "yes" to run expensive tests (default: no)
# PER_TEST_TIMEOUT Per-test timeout in seconds (default: 30)
# Don't exit on failure since test failures are expected.
set -o pipefail
# Configuration
RUST_GREP_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"
GNU_GREP_DIR="${GNU_GREP_DIR:-${RUST_GREP_DIR}/../gnu.grep}"
GNU_TESTS_DIR=""
VERBOSE=false
QUIET=false
JSON_OUTPUT_FILE=""
PER_TEST_TIMEOUT="${PER_TEST_TIMEOUT:-30}"
DETAILED_RESULTS=()
# Statistics
TOTAL_TESTS=0
PASSED_TESTS=0
FAILED_TESTS=0
SKIPPED_TESTS=0
usage() {
echo "Usage: $0 [options]"
echo
echo "Options:"
echo " -h, --help Show this help message"
echo " -v, --verbose Show diagnostics for failing/skipped tests"
echo " -q, --quiet Only print failures and the final summary"
echo " --json-output FILE Write results to FILE as JSON"
echo
echo "Environment variables:"
echo " GNU_GREP_DIR Path to the extracted GNU grep source tree"
echo " (default: ../gnu.grep)"
echo " RUN_EXPENSIVE_TESTS Set to 'yes' to run expensive tests"
echo " PER_TEST_TIMEOUT Per-test timeout in seconds (default: 30)"
echo
echo "Setup:"
echo " mkdir -p ../gnu.grep && (cd ../gnu.grep && bash ../grep/util/fetch-gnu.sh)"
}
log_info() { [[ "$QUIET" != "true" ]] && echo "[INFO] $1"; return 0; }
log_success() { [[ "$QUIET" != "true" ]] && echo "[PASS] $1"; return 0; }
log_skip() { [[ "$QUIET" != "true" ]] && echo "[SKIP] $1"; return 0; }
log_warning() { echo "[WARN] $1"; }
log_error() { echo "[FAIL] $1"; }
# Generate JSON output (schema shared with ../sed so compare_test_results.py works).
generate_json_output() {
cd "$RUST_GREP_DIR" || return
local timestamp
timestamp=$(date -u +"%Y-%m-%dT%H:%M:%SZ")
local rust_version
rust_version=$(cargo metadata --no-deps --format-version 1 2>/dev/null | jq -r '.packages[0].version // "unknown"')
local tests_json="[]"
if [[ ${#DETAILED_RESULTS[@]} -gt 0 ]]; then
local temp_file
temp_file=$(mktemp)
printf "%s\n" "${DETAILED_RESULTS[@]}" > "$temp_file"
tests_json=$(jq -s '.' < "$temp_file" 2>/dev/null) || tests_json="[]"
rm -f "$temp_file"
fi
jq -n \
--arg timestamp "$timestamp" \
--argjson total "$TOTAL_TESTS" \
--argjson passed "$PASSED_TESTS" \
--argjson failed "$FAILED_TESTS" \
--argjson skipped "$SKIPPED_TESTS" \
--argjson duration "$duration" \
--arg rust_version "$rust_version" \
--arg gnu_testsuite_dir "$GNU_TESTS_DIR" \
--argjson tests "$tests_json" \
'{
timestamp: $timestamp,
summary: {
total: $total,
passed: $passed,
failed: $failed,
skipped: $skipped,
duration_seconds: $duration
},
environment: {
rust_grep_version: $rust_version,
gnu_testsuite_dir: $gnu_testsuite_dir
},
tests: $tests
}' > "$JSON_OUTPUT_FILE"
log_info "JSON results written to: $JSON_OUTPUT_FILE"
}
# Parse command line arguments
while [[ $# -gt 0 ]]; do
case $1 in
-h|--help) usage; exit 0 ;;
-v|--verbose) VERBOSE=true; shift ;;
-q|--quiet) QUIET=true; shift ;;
--json-output) JSON_OUTPUT_FILE="$2"; shift 2 ;;
*) echo "Unknown argument: $1"; usage; exit 1 ;;
esac
done
# Validate environment
if [[ -d "$GNU_GREP_DIR" ]]; then
GNU_GREP_DIR="$(cd "$GNU_GREP_DIR" && pwd)"
GNU_TESTS_DIR="$GNU_GREP_DIR/tests"
fi
if [[ ! -f "$GNU_TESTS_DIR/init.sh" ]]; then
log_error "GNU grep testsuite not found at: $GNU_GREP_DIR"
log_error "Fetch it with:"
log_error " mkdir -p ${RUST_GREP_DIR}/../gnu.grep && (cd ${RUST_GREP_DIR}/../gnu.grep && bash ${RUST_GREP_DIR}/util/fetch-gnu.sh)"
exit 1
fi
if [[ ! -f "$RUST_GREP_DIR/Cargo.toml" ]]; then
log_error "Not in a Rust project directory: $RUST_GREP_DIR"
exit 1
fi
# Build the Rust grep implementation
log_info "Building Rust grep implementation..."
cd "$RUST_GREP_DIR" || exit 1
if ! cargo build --release --quiet; then
log_error "Failed to build Rust grep implementation"
exit 1
fi
RUST_GREP_BIN="$RUST_GREP_DIR/target/release/grep"
if [[ ! -x "$RUST_GREP_BIN" ]]; then
log_error "Built grep binary not found at: $RUST_GREP_BIN"
exit 1
fi
log_info "Using Rust grep binary: $RUST_GREP_BIN"
# Create a temporary work tree that mimics a GNU grep build directory.
TEST_WORK_DIR=$(mktemp -d)
trap 'rm -rf "$TEST_WORK_DIR"' EXIT
log_info "Test working directory: $TEST_WORK_DIR"
# A fake $abs_top_builddir whose src/ holds the binaries the tests expect.
BUILD_DIR="$TEST_WORK_DIR/build"
BIN_DIR="$BUILD_DIR/src"
mkdir -p "$BIN_DIR"
# grep, plus the egrep/fgrep wrappers a handful of tests rely on.
cat > "$BIN_DIR/grep" <<WRAPPER_EOF
#!/bin/sh
exec "$RUST_GREP_BIN" "\$@"
WRAPPER_EOF
cat > "$BIN_DIR/egrep" <<WRAPPER_EOF
#!/bin/sh
exec "$RUST_GREP_BIN" -E "\$@"
WRAPPER_EOF
cat > "$BIN_DIR/fgrep" <<WRAPPER_EOF
#!/bin/sh
exec "$RUST_GREP_BIN" -F "\$@"
WRAPPER_EOF
chmod +x "$BIN_DIR/grep" "$BIN_DIR/egrep" "$BIN_DIR/fgrep"
# Empty config.h: tests that probe it for build-time features just skip.
: > "$BUILD_DIR/config.h"
# get-mb-cur-max is a tiny standalone helper used by the locale require_ checks.
if [[ -f "$GNU_TESTS_DIR/get-mb-cur-max.c" ]]; then
if cc -I"$BUILD_DIR" -o "$BIN_DIR/get-mb-cur-max" "$GNU_TESTS_DIR/get-mb-cur-max.c" 2>/dev/null; then
log_info "Built get-mb-cur-max helper"
else
log_warning "Could not build get-mb-cur-max; multibyte/locale tests may skip"
fi
fi
# Replicate the PCRE_WORKS probe from tests/Makefile.am's TESTS_ENVIRONMENT.
PCRE_WORKS=0
if err=$(echo . | "$BIN_DIR/grep" -Pq . 2>&1); then
[[ -z "$err" ]] && PCRE_WORKS=1
fi
log_info "PCRE_WORKS=$PCRE_WORKS"
GREP_VERSION=$(basename "$GNU_GREP_DIR" | sed 's/^grep-//')
[[ "$GREP_VERSION" == "$(basename "$GNU_GREP_DIR")" ]] && GREP_VERSION="unknown"
HOST_TRIPLET="$(uname -m)-pc-linux-gnu"
# Record a test result (for JSON output)
record_result() {
if [[ -n "$JSON_OUTPUT_FILE" ]]; then
DETAILED_RESULTS+=("$(jq -n \
--arg name "$1" --arg status "$2" --arg error "$3" \
'{name: $name, status: $status, error: $error}')")
fi
}
# Run a single GNU testsuite script with the Rust grep on PATH.
run_gnu_test() {
local test_script="$1"
local test_name
test_name=$(basename "$test_script")
TOTAL_TESTS=$((TOTAL_TESTS + 1))
local test_output_file="$TEST_WORK_DIR/test_output_$$"
local test_exit_code=0
# When not the process-group leader (e.g. in CI), GNU timeout falls back to
# "foreground" mode and SIGTERMs the whole group on timeout. Shield the
# parent script so a single hung test doesn't take the run down.
trap '' TERM
(
cd "$TEST_WORK_DIR" || exit 99
# init.cfg refuses to run if these are set.
unset GREP_COLOR GREP_COLORS TERM CDPATH
export PATH="$BIN_DIR:$PATH"
export srcdir="$GNU_TESTS_DIR" abs_srcdir="$GNU_TESTS_DIR"
export abs_top_srcdir="$GNU_GREP_DIR" top_srcdir="$GNU_GREP_DIR"
export abs_top_builddir="$BUILD_DIR"
export CONFIG_HEADER="$BUILD_DIR/config.h"
export built_programs="grep egrep fgrep"
export AWK=awk PERL=perl SHELL=/bin/sh MAKE=make CC=cc
export LC_ALL=C MALLOC_PERTURB_=87
export VERSION="$GREP_VERSION" PACKAGE_VERSION="$GREP_VERSION"
export host_triplet="$HOST_TRIPLET"
export PCRE_WORKS="$PCRE_WORKS"
export GREP_TEST_NAME="$test_name"
export RUN_EXPENSIVE_TESTS="${RUN_EXPENSIVE_TESTS:-no}"
# fd 9 is the framework's stderr (init.cfg's stderr_fileno_=9).
if [[ "$test_name" == *.pl ]]; then
exec timeout --kill-after=5 "$PER_TEST_TIMEOUT" \
perl -w -I"$GNU_TESTS_DIR" -MCoreutils -MCuSkip "$test_script" 9>&2
else
exec timeout --kill-after=5 "$PER_TEST_TIMEOUT" \
/bin/sh "$test_script" 9>&2
fi
) </dev/null >"$test_output_file" 2>&1
test_exit_code=$?
trap - TERM
# Strip NUL bytes: some tests (e.g. z-anchor-newline) emit binary output,
# which would otherwise trigger a "ignored null byte" warning from $(...).
local test_output=""
[[ -f "$test_output_file" ]] && test_output=$(tr -d '\0' < "$test_output_file")
rm -f "$test_output_file"
# 124 = GNU timeout, 125 = uutils timeout, >=128 = killed by signal.
if [[ $test_exit_code -eq 124 || $test_exit_code -eq 125 || $test_exit_code -ge 128 ]]; then
log_error "$test_name (timeout)"
FAILED_TESTS=$((FAILED_TESTS + 1))
record_result "$test_name" "FAIL" "Test timed out after ${PER_TEST_TIMEOUT}s"
return
fi
case $test_exit_code in
0)
log_success "$test_name"
PASSED_TESTS=$((PASSED_TESTS + 1))
record_result "$test_name" "PASS" ""
;;
77)
log_skip "$test_name"
SKIPPED_TESTS=$((SKIPPED_TESTS + 1))
[[ "$VERBOSE" == "true" ]] && echo "$test_output" | head -3 | sed 's/^/ | /'
record_result "$test_name" "SKIP" "$test_output"
;;
*)
log_error "$test_name (exit $test_exit_code)"
FAILED_TESTS=$((FAILED_TESTS + 1))
[[ "$VERBOSE" == "true" ]] && echo "$test_output" | head -10 | sed 's/^/ | /'
record_result "$test_name" "FAIL" "Exit code: $test_exit_code"
;;
esac
}
# Discover the canonical test list from tests/Makefile.am's TESTS variable.
collect_tests() {
awk '
/^TESTS *\+?=/ { collect=1; sub(/^TESTS *\+?=/, "") }
collect {
line=$0
cont=sub(/\\[ \t]*$/, "", line)
n=split(line, a, /[ \t]+/)
for (i=1; i<=n; i++) if (a[i] != "") print a[i]
if (!cont) collect=0
}
' "$GNU_TESTS_DIR/Makefile.am"
}
log_info "Discovering tests from $GNU_TESTS_DIR/Makefile.am"
mapfile -t TEST_LIST < <(collect_tests | sort -u)
log_info "Found ${#TEST_LIST[@]} tests"
log_info "Starting test execution..."
start_time=$(date +%s)
for t in "${TEST_LIST[@]}"; do
[[ -z "$t" ]] && continue
test_path="$GNU_TESTS_DIR/$t"
[[ -f "$test_path" ]] || { log_warning "Listed test not found: $t"; continue; }
run_gnu_test "$test_path"
done
end_time=$(date +%s)
duration=$((end_time - start_time))
# Print summary
echo
echo "========================================="
echo "GNU grep testsuite results"
echo "========================================="
echo "Total tests: $TOTAL_TESTS"
echo "Passed: $PASSED_TESTS"
echo "Failed: $FAILED_TESTS"
echo "Skipped: $SKIPPED_TESTS"
echo "Duration: ${duration}s"
if [[ -n "$JSON_OUTPUT_FILE" ]]; then
generate_json_output
fi
if [[ $((PASSED_TESTS + FAILED_TESTS)) -gt 0 ]]; then
pass_rate=$(( (PASSED_TESTS * 100) / (PASSED_TESTS + FAILED_TESTS) ))
echo "Pass rate: ${pass_rate}%"
fi
# Mirror the script's exit convention to ../sed: nonzero if anything failed.
[[ $FAILED_TESTS -eq 0 ]] && exit 0 || exit 1