Data management¶
/admin is the administrator's view of the data behind the dashboard. It exists
so that whoever is responsible for the deployment can see exactly which files are
being served, take a copy of them, read what shape a replacement must be, and
review who has done what.
Replacing the data is not available yet
The page can show you the data and give you a copy of it. It cannot put new data live. The Replace buttons on the page are disabled, and the underlying endpoint, if called directly, validates and stores a workbook without publishing it. Nothing on this screen changes what the dashboard serves. See What is not built yet.
Who can reach it¶
There are exactly two roles.
| Role | Can do |
|---|---|
viewer |
Sign in, use the dashboard, comparison view, exports and reports |
admin |
All of the above, plus /admin and every /api/admin/* endpoint |
The role is resolved on the server at sign-in, from the verified Firebase token,
and stored in the session cookie, which is signed with SECRET_KEY. The browser
never sends a role and is never asked for one, so a user cannot promote
themselves by editing the cookie.
Resolution runs in this order, and the first match wins:
BYPASS_AUTH=1givesadmin. This is a local-development setting that switches off authentication for the whole process. Never set it on a deployment.- A Firebase custom claim
roleon the user's token, if its value names a role the application knows. This is the real mechanism: the claim is set out of band with the Firebase Admin SDK, travels with the identity, and cannot be edited by the user. - The
ADMIN_EMAILSallowlist, a comma-separated environment variable matched case-insensitively against the signed-in address. It matches only an email Firebase has marked verified. Firebase permits an account to hold an unverified address, so matching on one would let anyone who can sign up claim an administrator's email and be promoted on the spot. - Anything else gives
viewer. An unrecognised claim value, a missing claim, an empty allowlist and an unknown user all land here. No code path reachesadminby accident.
The allowlist lives in an environment variable rather than a table in the application, so the set of people who can overwrite the dataset changes only through a deployment.
If the deployment is configured with ADMIN_USERNAME and ADMIN_PASSWORD, that
break-glass login also grants admin. It exists only when both are set in the
environment, which is itself a deployment-level act.
/admin and every /api/admin/* route carry the check independently. The
navigation bar hides the Admin link from viewers, but that is presentation only:
hiding the link protects nothing and revealing it costs nothing.
The page also carries a banner it shows only when BYPASS_AUTH is active, saying
that the gate it is sitting behind is currently switched off, and whether the
flag came from the platform environment or from an app/.env file inside the
image.
Taking a copy of the current data¶
Download all current data produces swag-dss-data.zip containing the 13
served files, about 83 MB. The zip is assembled in memory, so nothing is left
behind on the server if the request fails.
The Files the dashboard is using panel lists each file with its live row count, its size and a per-file download link, so a single table can be taken without pulling the whole set.
Both routes write an entry to the audit log. That log is what the upload
endpoint checks before it will accept anything: a replacement is refused with
409 until a download has been recorded on this deployment, so there is always
something to go back to. The check is against the recorded fact rather than a
confirmation checkbox.
Source datasets and their formats¶
The dashboard's CSV and GeoJSON files are produced offline from four source datasets. The page shows the contract for each; they are reproduced here so they can be read before anyone starts preparing a file.
Model results workbook¶
Source id workbook-114. An Excel workbook, .xlsx. Every water-balance,
reallocation and storage number on the dashboard originates here, and it produces
reallocation.csv, sat_bands.csv, combined_bands.csv, simulation_sat.csv
and simulation_combined.csv.
It must contain all three of these sheets, named exactly:
| Sheet | Machine column names on | Required columns |
|---|---|---|
Reallocat_Ec85_Ea70 |
spreadsheet row 3 | 29, including hydro_id |
SAT_Ec85_Ea70_req |
spreadsheet row 2 | 35: 11 identity and storage columns, plus 12 *_vol and 12 *_area monthly columns |
Both_Ec85_Ea70_req |
spreadsheet row 2 | the same 35 |
The Reallocat sheet has three header rows (section, description, machine
names); the other two have two. Every sheet must have data rows below the header.
The monthly columns are named for the month with a two-digit prefix: 01Jan_vol
through 12Dec_vol, and 01Jan_area through 12Dec_area.
The sheet names encode the efficiency assumptions (Ec85_Ea70) as literal
strings. A model run under different assumptions needs a code change before it
can be read, not just a differently named sheet.
Sub-watershed polygons¶
Source id polygons-114. A GeoPackage, .gpkg, in EPSG:32637. It is the
delineation the workbook's rows refer to, and produces
elevation_bands.geojson and flow_edges.geojson.
band_id in the polygons is the workbook's hydro_id. The two files are one
dataset and must be replaced together. Replacing one alone makes the map joins
drop features silently, which looks like missing data rather than like an error.
WRUA boundaries¶
Source id boundaries-chef. A shapefile in a ZIP, 119 polygons. All four
companion files (.shp, .dbf, .shx, .prj) must be present. It produces
wrua_boundaries.geojson and other_wruas.geojson.
Attribute changes and boundary corrections are safe on their own. Adding or removing a WRUA is a model change and needs a matching workbook.
Agro-ecological crop zones¶
Source id crop-zones. A shapefile, major_crop.shp. The Major_Crop
attribute drives the dissolve. It produces crop_zones.geojson and
crop_recommendations.csv. This is reference data and is effectively static.
What the validator checks¶
POST /api/admin/upload/<source_id> runs the model results workbook through a
validator before storing it. The validator answers one question: would the
offline ingest succeed on this file?
It reports every problem it finds in one response, naming the sheet and the columns, rather than stopping at the first. A large file over a slow connection should not have to be uploaded once per fault.
It checks that the file opens as a workbook, that all three sheets are present, that each sheet carries its required columns on the expected header row, and that each has data rows. When a column is missing it says which spreadsheet row the names were read from, because headings on the wrong row is the usual cause.
It does not run the ingest and it does not check the numbers. A workbook can satisfy every rule here and still contain nonsense. This is a gate against the file being the wrong shape, not against the model being wrong.
Alongside the result it returns a summary: file size, the sheets it found, row and column counts per sheet, and the number of distinct WRUA names, so an administrator can confirm the file is the one they meant.
Activity log¶
The Activity panel shows the most recent entries, newest first, each with who acted, what they did and when. Downloads, staged uploads and rejected uploads are all recorded.
The log is append-only JSON Lines in the runtime directory, which defaults to
app/data/uploads/ and can be pointed elsewhere with RUNTIME_DIR. On a host
with an ephemeral filesystem, Render among them, that directory does not survive
a redeploy, so the log covers the current deployment and no further. Nothing else
depends on it: the application works either way, it just forgets who downloaded
what.
What is not built yet¶
Read this before promising anyone that data can be updated through the browser.
- The page cannot upload. Every Replace button is rendered disabled, and the page shows the reason the inventory reports. It is shown rather than hidden so an administrator is not left wondering whether they are on the wrong screen.
- The endpoint does not publish.
POST /api/admin/upload/<source_id>validates a workbook and writes it into a staging directory. Its response says"published": falseand states in words that the file is not live. Nothing reads the staged file afterwards. - No pipeline runs on the server. The step that turns a source workbook into
the served CSV and GeoJSON files runs offline, and
app/data_prepis excluded from the container image, so there is nothing on the server that could process an upload even if one were accepted. - Nothing triggers a reload. The application does have the machinery to pick
up changed files:
reload_data()re-reads everything and clears the GeoJSON cache, and a stamp file lets one worker tell the others to do the same. No code path calls it, because no code path publishes.
Until those are in place, replacing the data means preparing the files offline and deploying them, and this page is where you get the current copy and the format contract to work against.