Developer Documentation
Welcome to the Developer Documentation for the South Dakota Rain and Drought Analysis project. This guide provides technical onboarding, architectural details, data pipeline schema description, and local build instructions.
1. Project Directory & Architecture
This repository is structured as a Quarto Website that processes and visualizes 10 years of South Dakota mean daily precipitation and county-level weekly U.S. Drought Monitor (USDM) statistics.
Key Directories & Files
rainDrought/: A dedicated Python package housing the core climate data collection, cleaning, and visualization scripts:config.py: Dynamic parser for regional county/state custom settings.data_pipeline.py: Automates REST API fetching from USDM and ACIS endpoints.visualizations.py: Handles data preprocessing, Matplotlib charts, and interactive Plotly subplots.
_quarto.yml: Quarto configuration setting format options, navbar layout, and build output paths.index.qmd: Welcome page containing overview texts and data source links.collect_data.qmd: Page compiling the data fetching pipeline (calls functions inrainDrought.data_pipeline).plot_records.qmd: Page compiling interactive Plotly dashboards (calls functions inrainDrought.visualizations).data/: Cache folder containing CSV datasets for precipitation and county drought records.
Output Layout
The site is compiled into the docs/ directory for GitHub Pages: - docs/index.html (rendered overview) - docs/collect_data.html (rendered data collection logs) - docs/plot_records.html (rendered interactive Plotly visualizations) - docs/search.json (search index for site-wide navigation)
[!IMPORTANT] The directory
docs/site_libs/is ignored via.gitignorebecause all pages are compiled as standalone HTML documents with assets embedded natively. Do not stage or commit files insidedocs/site_libs/.
2. Local Environment Setup
The workspace is configured to use a specific Python Conda environment.
Target Executable
Always run commands and scripts using the Earth Analytics Python environment:
/users/brianyandell/miniconda3/envs/earth-analytics-python/bin/pythonDependencies
The Python script blocks require the following libraries: - pandas (Data manipulation) - numpy (Vector operations) - requests (API requests) - plotly (Interactive visualizations) - matplotlib & seaborn (Static data validation charts)
To install or update dependencies in the environment, run:
/users/brianyandell/miniconda3/envs/earth-analytics-python/bin/python -m pip install pandas requests plotly matplotlib seaborn2.5. Location Configuration (config.csv)
By default, the project processes data for South Dakota, Oglala Lakota County, and Todd County. You can customize the state and counties analyzed by creating a config.csv file in the root of the repository.
Configuration Templates
The repository includes two pre-configured CSV templates: - config.csv.default: The default configuration (South Dakota, Pine Ridge, Rosebud) used for the live site. - config.csv.example: An example configuration (Nebraska, Lancaster County) showing custom region setups.
To set up a custom analysis, copy either of these templates to a new config.csv file in the root directory:
cp config.csv.default config.csv
# or
cp config.csv.example config.csvSchema Parameters
type: Row type, eitherstateorcounty.code: The lowercase state abbreviation (e.g.ne) or the FIPS county code (e.g.31107).name: Full official name of the state or county.label: Short label for plots and legends (e.g.RosebudorLancaster).
Output Naming Convention
- Default state:
data/south_dakota_precipitation_daily.csv - Custom state:
data/{state_name_lower}_precipitation_daily.csv - Default counties:
data/oglala_lakota_drought_weekly.csv,data/todd_drought_weekly.csv - Custom counties:
data/{county_name_lower}_drought_weekly.csv
The file config.csv is ignored in Git to prevent local region settings from overriding the default South Dakota pages deployed to GitHub Pages.
3. Data Pipeline & Endpoints
Data collection is automated within collect_data.qmd. When executed, the pipeline performs two main API tasks:
Task 1: Weekly U.S. Drought Monitor (USDM) County Data
- Source: National Drought Mitigation Center (NDMC)
- API Endpoint:
https://usdmdataservices.unl.edu/api/CountyStatistics/GetDroughtSeverityStatisticsByAreaPercent - Request Parameters:
aoi: FIPS county code (46102for Oglala Lakota County,46121for Todd County)startdate: Dynamic start date from 10 years ago (YYYY-MM-DD)enddate: Today’s date (YYYY-MM-DD)statisticsType:1(representing area percent)
- Outputs:
Task 2: Daily Precipitation (ACIS GridData)
- Source: NOAA Regional Climate Centers (RCCs)
- API Endpoint:
https://data.rcc-acis.org/GridData - JSON Payload Parameters:
state:"sd"(South Dakota)sdate/edate: Start and end date rangegrid:"1"(NRCC Interpolated Grid)elems:[{"name": "pcpn", "area_reduce": "state_mean"}]
- Output:
4. Data Processing Standards & Formulas
When fetching and cleaning raw data, follow these standards:
Drought Severity and Coverage Index (DSCI)
The weekly county-level USDM statistics include five severity categories (D0 to D4), representing the percentage area of the county experiencing that level of drought (or worse). The DSCI is calculated as the sum of these percentages: \[\text{DSCI} = \text{D0} + \text{D1} + \text{D2} + \text{D3} + \text{D4}\] - The resulting DSCI ranges between 0 (no drought) and 500 (100% of the county is in D4 Exceptional Drought).
Precipitation Clean-Up
- Trace Amounts: Daily precipitation data containing
'T'or'T 'are parsed and converted to0.0001inches. - Missing Data: Entries containing
'M'or empty strings are set toNone/NaNin the data frame.
Cumulative Aggregations
To evaluate annual progressions: - Group records by year (date.dt.year). - Compute cumulative sums (cumsum()) for precipitation and DSCI values starting January 1st of each year.
5. Visualizations & Plotly Controls
The dashboard page plot_records.qmd compiles three primary interactive dashboards:
- 10-Year Timelines: Displays weekly DSCI for both counties and the 365-day rolling statewide precipitation.
- Annual Progression over the Year: Standardizes month on the X-axis (
tickformat="%b") and cumulative values on the Y-axis. Open circles mark the end of each month. - Annual Trajectories of Rain vs. DSCI: Facets the cumulative rain (X-axis) vs. cumulative DSCI (Y-axis) side-by-side.
- Overlaid watermark annotations (
DRYin top-left,WETin bottom-right) indicate climate zones.
- Overlaid watermark annotations (
Plotly Interactive Legend Toggling
To allow simultaneous visibility toggling of a specific year across multiple subplots: - Curves share the same legendgroup name (e.g. legendgroup="2020"). - The parent layout has groupclick="togglegroup" set under the legend configuration.
6. Local Build & Deployment
Build the Entire Site
To compile the website assets locally, execute:
quarto renderThis command compiles index.qmd, collect_data.qmd, and plot_records.qmd into their corresponding HTML pages under the docs/ folder.
Standalone Execution Script
For non-Quarto workflows, you can execute the entire collection and plotting process as a standalone Python script:
python rainDrought/run_analysis.pyThis script: 1. Loads location configurations from config.csv (or falls back to default settings). 2. Automates API data collection and caches datasets to the data/ folder. 3. Prints dataset statistics and records directly to the command line. 4. Generates interactive Plotly plots and saves them as self-contained HTML dashboards in the output/ folder (output/time_series.html, output/cumulative_progression.html, and output/trajectories.html), then opens them in your default web browser.
Run Data Updates via Quarto
To run the pipeline and update the data caches using Quarto:
quarto render collect_data.qmdGitHub Pages Deployment
- Commit all modified source files (
.qmd,_quarto.yml,data/*.csv,DEVELOPER.md) and the compiled HTML outputs in/docs. - Do not commit files inside
docs/site_libs/. - Push the changes to GitHub.
- Set the GitHub Pages source branch to
main(or your active branch) and select/docsas the source folder in your repository settings.