bolivia112 is the province-level (112 provinces) replication of
DS4Bolivia (339 municipalities). Every municipal dataset has been aggregated to
Bolivia's 112 provinces using population-weighted aggregation. The folder structure, filenames
and column schemas mirror ds4bolivia; the primary join key is prov_id (the first 3 digits of the
INE mun_id).
π How the data was built: see
province_aggregation_report.md. Intensive variables (indices, rates, per-capita, embeddings) are population-weighted means; extensive variables (population, COβ, land-cover/area counts) are summed. All variables are weighted by the INEpop/pop.csvseries β SDG variables bypop2020, others year-matched β and thepopulation_2020column equals INE pop2020. Province cells with no municipal data for an SDG variable are imputed from the department's pop-weighted mean and flagged (<var>_imputed). Everything is reproducible viacode/build_bolivia112.py.
This repository is organized for researchers and data scientists interested in:
- Spatial Econometrics: Understanding regional disparities, growth, and clustering.
- Spatial Machine Learning: Utilizing satellite imagery (Earth Observation) for predictive modeling.
- Sustainable Development: Tracking SDG indicators at the province level.
Operational β every 112-province dataset is complete and the build is fully reproducible
(uv run python code/build_bolivia112.py, idempotent). Last updated: 2026-06-23.
| Area | Status |
|---|---|
| Data | β
Complete β 112 provinces in every folder; master bolivia112_v20260622.csv (112 Γ 352) |
| Build pipeline | β
Reproducible & idempotent β per-variable rules in code/aggregation_rules.csv |
| SDG aggregation | β
Audited β weighted by INE pop2020, areal rates area-weighted, all-NaN cells imputed + flagged |
| Maps | β 112 province boundaries (GADM ADM2) |
| Notebooks | β 6 province-level EDA/ESDA tutorials |
| Web app |
Recent changes (last 2 commits): SDG weighting switched from the Municipal-Atlas count to INE
pop2020; the three areal rates re-weighted by land area; 8 all-NaN SDG cells imputed + flagged; and a
full statistical audit added (sdg_aggregation_audit.md).
Before using the data (see sdg_aggregation_audit.md for detail):
- 36 of 62 SDG indicators are population-weighted approximations of non-population rates β check
sdgVariables/sdg_reference_populations.csv(weight_is_approximate) andcode/sdg_aggregation_sensitivity.csv. - Exclude
_imputedcells from inference; readimds/index_sdg*as averaged municipal indices (not province-recomputed); MAUP applies and 15 provinces are single-municipality passthroughs.
Pending: the Earth Engine app needs an uploaded province boundary asset to run live.
- Space-time dynamics of population, luminosity, land cover and GDP: Google Earth Engine GeoExplorer, adapted to province boundaries. (Requires a province boundary asset in your Earth Engine account β see apps/README.md.)
Province-level adaptations of the DS4Bolivia tutorials (GeoPandas, PySAL, scikit-learn). Each
notebook now groups by province (ADM2 = 'prov') over the 112 provinces.
- EDA β descriptive statistics, regional comparisons, NTL vs development.
- ESDA β spatial clusters and outliers via Global/Local Moran's I.
- Spatial Distribution & Dependence β classification schemes, spatial weights, LISA clusters.
- Spatial Inequality β Theil/Gini decomposition by department.
- Spatial Heterogeneity (GWR & MGWR) β spatially varying NTLβdevelopment relationships.
- Extended EDA + Spatial Analysis β combined traditional + spatial methods.
- GDP Validation β space-time dynamics of GDP per capita (1990β2024) and a check that the aggregated SDG indices/indicators correlate with GDP in the theoretically expected directions.
See notebooks/README.md. Sources are kept as Jupytext .md; regenerate
executable notebooks with uv run jupytext --to ipynb --execute notebooks/<name>.md.
All province datasets use prov_id as the primary join key.
| Dataset | Description | Documentation |
|---|---|---|
| regionNames | Administrative metadata for 112 provinces (name, capital, department, n_mun) | README |
| sdg | Aggregated SDG indices (population-weighted) | README |
| sdgVariables | 64 granular SDG indicators (population-weighted) | README |
| pop | Population time series 2001-2020 (summed) | README |
| ntl | Night-time lights, ln per capita 2012-2020 (population-weighted) | README |
| satelliteEmbeddings | 64-dim embeddings, 2017 (population-weighted) | README |
| datasets | Pre-merged SDGs + Satellite Embeddings | README |
| gdp | Province GDP per capita 1990-2024 (GADM ADM2) | README |
| Resource | Description |
|---|---|
| maps | 112 province boundaries (GeoJSON), from the GADM ADM2 GeoPackage |
| Resource | Description | Documentation |
|---|---|---|
| code | build_bolivia112.py (reproducible aggregation) + aggregation_rules.csv + adapted ML models |
README |
| notebooks | Province-level ESDA / spatial analysis tutorials | README |
| apps | Interactive GeoExplorer (province) | README |
Each folder's README carries the full detail; this is the one-line summary. "Aggregation" means
municipality β province (339 β 112) β see province_aggregation_report.md
and the per-variable table code/aggregation_rules.csv.
| Dataset | Aggregation (rule Β· weight) | Generated from (original source) |
|---|---|---|
regionNames |
not aggregated β derived identifiers | INE mun_id + code/provinceNames.csv (GADM/INE/Wikipedia cross-checked) |
sdg, sdgVariables |
population-weighted mean Β· pop2020; all-NaN cells imputed (dept mean) + flagged |
Atlas Municipal de los ODS β Andersen et al. (2020), SDSN Bolivia |
pop |
sum (extensive count) | code/archive_stata_code/040_*.do from a raw export β provider not documented in repo |
ntl |
population-weighted mean of ln values Β· year-matched pop |
VIIRS night-time lights β 030_*.do / 060_*.do (HP trend Ξ»=400) |
satelliteEmbeddings |
population-weighted mean Β· pop2017 (pop_sum summed in the popWeighted file) |
Google Satellite Embeddings V1 (2017) via GEE β code/aggregate-satellite-embedings-to-adm*.js |
datasets |
join of sdg + satelliteEmbeddings (rules inherited) |
as above |
gdp |
not aggregated β attached as-is | GADM ADM2 GeoPackage (already province-level) β original GDP method not documented in repo |
maps |
not aggregated β geometry | GADM ADM2 GeoPackage; polygons matched to prov_id by name |
master bolivia112_v*.csv |
per-variable via classify() β see below |
all of the above, from ds4bolivia |
bolivia112_v20260622.csv gathers every municipal variable in one wide table. Each variable's
aggregation rule is recorded per-variable in
code/aggregation_rules.csv (varname, agg, weight, note):
wmean(139) β intensive variables: allsdg*/index_sdg*/imds/urbano_2012weighted bypop2020;ln_NTLpc*, embeddings, climate/elevation/distance.mean, MODIS ratios weighted by year-matchedpop.sum(138) β extensive totals:pop20YY,co20YY, COβ/land-cover/impervious-area pixel counts (population_2020is set to INEpop2020).min/max(32/32) β.min/.maxcompanions of physical variables.recompute(1) βrank_imds, re-ranked over the 112 provinces.- All-NaN SDG cells are imputed (department pop-weighted mean) and flagged via
<var>_imputed.
| File | Description |
|---|---|
bolivia112_v20260622.csv |
All 342 aggregated variables (+5 <var>_imputed flags) in one wide table (112 Γ 352) |
definitions_bolivia112_v20260622.csv |
Variable dictionary (one row per column) |
Mendez, C., Gonzales, E., Leoni, P., Andersen, L., Peralta, H. (2026). bolivia112: A Province-Level Data Science Repository to Study Regional Development in Bolivia [Data set]. GitHub. https://github.com/quarcs-lab/bolivia112
@misc{bolivia112_2026,
author = {Mendez, Carlos and Gonzales, Erick and Leoni, Pedro and Andersen, Lykke and Peralta, Hendrix},
title = {{bolivia112}: A Province-Level Data Science Repository to Study Regional Development in Bolivia},
year = {2026},
publisher = {GitHub},
journal = {GitHub repository},
howpublished = {\url{https://github.com/quarcs-lab/bolivia112}}
}This project uses UV for Python package management.
git clone https://github.com/quarcs-lab/bolivia112.git
cd bolivia112
uv sync
uv run jupyter notebook # run notebooks
uv run python code/build_bolivia112.py # rebuild all province data from ds4boliviaAll modules link by the unique province identifier prov_id.
| Dataset Category | File Path | Join Key |
|---|---|---|
| Region Names | regionNames/regionNames.csv |
prov_id |
| Socio-Economic | sdg/sdg.csv |
prov_id |
| Detailed SDG | sdgVariables/sdgVariables.csv |
prov_id |
| Population | pop/pop.csv |
prov_id |
| Night-time Lights | ntl/ln_NTLpc.csv |
prov_id |
| Satellite Features | satelliteEmbeddings/satelliteEmbeddings2017.csv |
prov_id |
| Spatial Vector | maps/bolivia112provincesOpt.geojson |
prov_id |
| Pre-merged | datasets/sdgs_satelliteEmbeddings2017.csv |
prov_id |
β οΈ Identifier note: the primary key for joining all province datasets isprov_id(3-digit INE province code, e.g.405= Litoral). Treat it as anintconsistently across dataframes before merging.
import pandas as pd
import geopandas as gpd
import matplotlib.pyplot as plt
BASE = "bolivia112" # local path, or a raw GitHub URL once published
df_names = pd.read_csv(f"{BASE}/regionNames/regionNames.csv")
df_sdg = pd.read_csv(f"{BASE}/sdg/sdg.csv")
df_emb = pd.read_csv(f"{BASE}/satelliteEmbeddings/satelliteEmbeddings2017.csv")
df = df_names.merge(df_sdg, on="prov_id").merge(df_emb, on="prov_id")
print(f"{len(df)} provinces, {len(df.columns)} columns")
gdf = gpd.read_file(f"{BASE}/maps/bolivia112provincesOpt.geojson")
gdf["prov_id"] = gdf["prov_id"].astype(int)
gdf = gdf.merge(df, on="prov_id", how="inner")
fig, ax = plt.subplots(figsize=(12, 10))
gdf.plot(column="index_sdg1", cmap="viridis", linewidth=0.1, edgecolor="white",
legend=True, legend_kwds={"label": "SDG 1 Index (No Poverty)", "orientation": "horizontal"}, ax=ax)
ax.set_title("Bolivia: SDG 1 Index by Province (112 provinces)", fontsize=15)
ax.set_axis_off()
plt.show()Find an error? Have a suggestion? Submit an issue or join the discussion via GitHub.
