A high-throughput Apache CLF log processing pipeline that parses and aggregates server logs significantly faster than pandas, using Numba for parallel byte-level parsing and JAX for vectorised aggregation and anomaly detection.
Pandas-based log analysis pipelines become the bottleneck at scale — a 1 GB log file (≈ 12.6 million lines) takes nearly 5 minutes to parse and aggregate with pandas on an 8-core machine. This project replaces that workflow with a compiled, parallel pipeline that parses the same file in ~2.4 seconds and runs the full parse + aggregate + anomaly-detection pass in ~11 seconds (26× faster end-to-end).
Apache CLF log file (.log)
│
▼
┌───────────────────────────────┐
│ src/parser.py │ Numba @njit(parallel=True)
│ Byte-level parallel parser │ Reads raw uint8 buffer, dispatches
│ │ one thread per log line via prange
└───────────────┬───────────────┘
│ NumPy arrays
│ (timestamps, status_codes,
│ response_sizes, ip_ints,
│ endpoint_hashes)
▼
┌───────────────────────────────┐
│ src/aggregator.py │ JAX @jit — single XLA pass
│ Vectorised aggregation │ RPM · error rate · p50/p95/p99
│ │ · top-10 endpoints
└───────────────┬───────────────┘
│ JAX arrays
▼
┌───────────────────────────────┐
│ src/anomaly.py │ JAX lax.scan — rolling z-score
│ Anomaly detector │ Flags windows where |z| > 3σ
│ │ Returns timestamps + severity
└───────────────┬───────────────┘
│
▼
Summary report + results/benchmark.json
Measured on 2026-10-03 — 8-core Linux machine, 31 GB RAM, JAX CPU backend (numba 0.68, jax 0.10.2, pandas 3.0.6). Times are end-to-end after JIT warm-up. Re-run colab_benchmark.ipynb (or python -m benchmark.bench_full) for your own hardware.
| Engine | File | Lines | Time (s) | Lines/s | Peak Mem | Speed-up vs pandas |
|---|---|---|---|---|---|---|
| pandas | 1 GB | 12,650,000 | 283.96 | 44,549 | 5.0 GB | 1.0× |
| numba (parse only) | 1 GB | 12,650,000 | 2.39 | 5,296,136 | 2.1 GB | 118.9× |
| numba + jax (cpu) | 1 GB | 12,650,000 | 10.82 | 1,169,204 | 2.1 GB | 26.2× |
Click Open in Colab above. The notebook installs dependencies, generates a 1 GB log file, runs all benchmarks, and downloads benchmark.json with the results.
git clone https://github.com/pradneshfernandez/log-analytics-engine
cd log-analytics-engine
pip install -r requirements.txt
# Generate a 1 GB test file
python data/generate_logs.py --output data/1gb.log --size-gb 1
# Run the full pipeline
python -c "from src.pipeline import run; run('data/1gb.log')"log-analytics-engine/
├── data/
│ └── generate_logs.py # Synthetic Apache CLF generator with anomaly injection
├── src/
│ ├── parser.py # Numba parallel byte-level parser + pure-Python fallback
│ ├── aggregator.py # JAX single-pass aggregation (RPM, error rate, percentiles)
│ ├── anomaly.py # JAX rolling z-score anomaly detector
│ └── pipeline.py # Orchestrator — wires all steps, prints summary report
├── benchmark/
│ ├── bench_pandas.py # Pandas baseline
│ ├── bench_numba.py # Parser-only benchmark
│ └── bench_full.py # Full comparison table (CLI)
├── colab_benchmark.ipynb # End-to-end Colab notebook
└── requirements.txt
Apache CLF parsing is embarrassingly parallel — each line is independent. Numba's @njit(parallel=True) compiles the parser to native machine code and distributes lines across CPU threads via prange, bypassing Python's GIL entirely. The parser works directly on a uint8 byte buffer rather than Python strings, which eliminates object allocation overhead.
Every Numba kernel has a pure-Python equivalent (parse_line_python) for unit testing without requiring compilation.
JAX's @jit decorator compiles the aggregation to an XLA program that runs all metrics — requests per minute, error rates, latency percentiles, endpoint counts — in a single fused pass. On a GPU runtime (Colab), the same code runs on the GPU with zero changes. jax.lax.scan is used for the rolling z-score so the anomaly detector compiles to a single XLA loop rather than a Python loop over time windows.
| Package | Role |
|---|---|
| numba | Parallel log line parser |
| jax[cpu] | Vectorised aggregation + anomaly detection |
| numpy | Array I/O between stages |
| pandas | Baseline benchmark only |
MIT