Senior Data & AI Engineer building governed AI systems, reliable data pipelines, and automation for real workflows.
I design and implement Python/SQL systems that turn fragmented data, documents, and operational processes into usable infrastructure: MCP and context gateways, retrieval and document pipelines, reproducible ML experiments, workflow automation, analytical products, and public-facing decision surfaces.
I usually work across the full path:
messy problem
→ explicit boundaries and contracts
→ working software and data pipelines
→ validation, evidence, and observability
→ documentation and human handoff
My background combines engineering practice with a PhD in Economics and an MSc in Physics. I am strongest where software, data, research judgment, and operational usefulness need to meet.
Main site · CV · Developer journal · Stack I actually use
A local, read-only MCP server and client that gives AI systems governed access to four independent knowledge sources without exposing arbitrary filesystem access.
The project demonstrates logical resource design, capability negotiation, explicit mappings, path-containment and symlink defenses, bounded reads, provenance-rich responses, structured failures, tests, an acceptance probe, and real client/server use over stdio.
A live news-intelligence and editorial pipeline organized around a practical operating path:
news acquisition → structured brief → draft → human last mile
It combines staged data processing, AI-assisted editorial workflows, public snapshots, explicit handoff surfaces, smoke checks, and a deployed publication interface.
A Python/SQL accounting spine that converts messy ledger inputs into canonical records, materialized views, metrics, drill-downs, and human-readable reports.
The system emphasizes explicit stages, reproducible commands, structured logging, data-quality diagnostics, professional handoff artifacts, and a clear boundary between operational truth and presentation.
A reproducible machine-learning experiment layer for income prediction using Argentina's Permanent Household Survey.
The repository separates upstream preprocessing from modeling responsibility and owns the target and feature contracts, leakage exclusions, split registry, guarded training runs, metrics, and thesis-ready evidence.
A set of independently useful repositories connected through explicit contracts rather than one shared internal architecture:
- KB Contracts — shared schemas, manifests, run records, stable IDs, publishing contracts, and integration rules;
- Knowledge Inspect — ingestion and analysis seams for chat and paper-derived knowledge artifacts;
- KB Artifacts — deterministic, read-only evidence selection with provenance and manifests.
The MCP gateway is the narrow access layer over this ecosystem; the source systems remain authoritative.
A research data-engineering and collaboration surface built around a recovered archive of survey, conflict, service-delivery, public-works, and spatial data.
It maps datasets, notebooks, reusable outputs, validation state, and the path from archive recovery to renewed analysis without rewriting or obscuring the original evidence.
- Small, reliable components over invented universal platforms.
- Logical identities and provenance over hidden path knowledge.
- Bounded access and explicit authority over convenient overexposure.
- Contracts, run records, checks, and evidence over undocumented success.
- Human-reviewable handoffs over opaque automation.
- Incremental delivery with honest scope and clearly stated non-goals.
My core working stack includes Python, SQL, Bash, pandas, scikit-learn, REST APIs, LLM APIs, MCP, RAG, LangChain, ChromaDB, JSON Schema, SQLite/PostgreSQL, BigQuery, FastAPI/Flask, Docker, GitHub Actions, Next.js, Docusaurus, and Vercel.
The evidence-backed catalog is here:
Awesome Automation I Actually Use — 80 tools, formats, platforms, and operating principles tied to public repositories and field notes.
Research is part of my foundation, not the main organization of this profile.
- PhD thesis portal — long-form proof of my ability to study a difficult problem deeply, develop original arguments, and sustain a multi-year analytical programme.
- Concentration Is Not Scaling — a current reproducibility repository with formal results, deterministic demonstrations, identity tests, manuscript builds, and release artifacts.
- Earlier economic-geography tools: geo-correlation-structures, pRCA, and spatial-coexistence-metrics.
My academic path includes research experience at the Harvard Kennedy School Growth Lab and Caltech, alongside applied work in economics, public data, and spatial analysis.
I teach computer science and data-oriented university courses at the University of Buenos Aires.
- Evaluar App — an extensible platform for exercises, student questions, tutoring workflows, and AI-assisted feedback.
- Computational Linear Algebra — notebooks, code, datasets, and laboratory material connecting numerical linear algebra with practical computation.
Buenos Aires, Argentina · English and Spanish



