Skip to content

Data management

/admin is the administrator's view of the data behind the dashboard. It exists so that whoever is responsible for the deployment can see exactly which files are being served, take a copy of them, read what shape a replacement must be, and review who has done what.

Replacing the data is not available yet

The page can show you the data and give you a copy of it. It cannot put new data live. The Replace buttons on the page are disabled, and the underlying endpoint, if called directly, validates and stores a workbook without publishing it. Nothing on this screen changes what the dashboard serves. See What is not built yet.

Who can reach it

There are exactly two roles.

Role Can do
viewer Sign in, use the dashboard, comparison view, exports and reports
admin All of the above, plus /admin and every /api/admin/* endpoint

The role is resolved on the server at sign-in, from the verified Firebase token, and stored in the session cookie, which is signed with SECRET_KEY. The browser never sends a role and is never asked for one, so a user cannot promote themselves by editing the cookie.

Resolution runs in this order, and the first match wins:

  1. BYPASS_AUTH=1 gives admin. This is a local-development setting that switches off authentication for the whole process. Never set it on a deployment.
  2. A Firebase custom claim role on the user's token, if its value names a role the application knows. This is the real mechanism: the claim is set out of band with the Firebase Admin SDK, travels with the identity, and cannot be edited by the user.
  3. The ADMIN_EMAILS allowlist, a comma-separated environment variable matched case-insensitively against the signed-in address. It matches only an email Firebase has marked verified. Firebase permits an account to hold an unverified address, so matching on one would let anyone who can sign up claim an administrator's email and be promoted on the spot.
  4. Anything else gives viewer. An unrecognised claim value, a missing claim, an empty allowlist and an unknown user all land here. No code path reaches admin by accident.

The allowlist lives in an environment variable rather than a table in the application, so the set of people who can overwrite the dataset changes only through a deployment.

If the deployment is configured with ADMIN_USERNAME and ADMIN_PASSWORD, that break-glass login also grants admin. It exists only when both are set in the environment, which is itself a deployment-level act.

/admin and every /api/admin/* route carry the check independently. The navigation bar hides the Admin link from viewers, but that is presentation only: hiding the link protects nothing and revealing it costs nothing.

The page also carries a banner it shows only when BYPASS_AUTH is active, saying that the gate it is sitting behind is currently switched off, and whether the flag came from the platform environment or from an app/.env file inside the image.

Taking a copy of the current data

Download all current data produces swag-dss-data.zip containing the 13 served files, about 83 MB. The zip is assembled in memory, so nothing is left behind on the server if the request fails.

The Files the dashboard is using panel lists each file with its live row count, its size and a per-file download link, so a single table can be taken without pulling the whole set.

Both routes write an entry to the audit log. That log is what the upload endpoint checks before it will accept anything: a replacement is refused with 409 until a download has been recorded on this deployment, so there is always something to go back to. The check is against the recorded fact rather than a confirmation checkbox.

Source datasets and their formats

The dashboard's CSV and GeoJSON files are produced offline from four source datasets. The page shows the contract for each; they are reproduced here so they can be read before anyone starts preparing a file.

Model results workbook

Source id workbook-114. An Excel workbook, .xlsx. Every water-balance, reallocation and storage number on the dashboard originates here, and it produces reallocation.csv, sat_bands.csv, combined_bands.csv, simulation_sat.csv and simulation_combined.csv.

It must contain all three of these sheets, named exactly:

Sheet Machine column names on Required columns
Reallocat_Ec85_Ea70 spreadsheet row 3 29, including hydro_id
SAT_Ec85_Ea70_req spreadsheet row 2 35: 11 identity and storage columns, plus 12 *_vol and 12 *_area monthly columns
Both_Ec85_Ea70_req spreadsheet row 2 the same 35

The Reallocat sheet has three header rows (section, description, machine names); the other two have two. Every sheet must have data rows below the header.

The monthly columns are named for the month with a two-digit prefix: 01Jan_vol through 12Dec_vol, and 01Jan_area through 12Dec_area.

The sheet names encode the efficiency assumptions (Ec85_Ea70) as literal strings. A model run under different assumptions needs a code change before it can be read, not just a differently named sheet.

Sub-watershed polygons

Source id polygons-114. A GeoPackage, .gpkg, in EPSG:32637. It is the delineation the workbook's rows refer to, and produces elevation_bands.geojson and flow_edges.geojson.

band_id in the polygons is the workbook's hydro_id. The two files are one dataset and must be replaced together. Replacing one alone makes the map joins drop features silently, which looks like missing data rather than like an error.

WRUA boundaries

Source id boundaries-chef. A shapefile in a ZIP, 119 polygons. All four companion files (.shp, .dbf, .shx, .prj) must be present. It produces wrua_boundaries.geojson and other_wruas.geojson.

Attribute changes and boundary corrections are safe on their own. Adding or removing a WRUA is a model change and needs a matching workbook.

Agro-ecological crop zones

Source id crop-zones. A shapefile, major_crop.shp. The Major_Crop attribute drives the dissolve. It produces crop_zones.geojson and crop_recommendations.csv. This is reference data and is effectively static.

What the validator checks

POST /api/admin/upload/<source_id> runs the model results workbook through a validator before storing it. The validator answers one question: would the offline ingest succeed on this file?

It reports every problem it finds in one response, naming the sheet and the columns, rather than stopping at the first. A large file over a slow connection should not have to be uploaded once per fault.

It checks that the file opens as a workbook, that all three sheets are present, that each sheet carries its required columns on the expected header row, and that each has data rows. When a column is missing it says which spreadsheet row the names were read from, because headings on the wrong row is the usual cause.

It does not run the ingest and it does not check the numbers. A workbook can satisfy every rule here and still contain nonsense. This is a gate against the file being the wrong shape, not against the model being wrong.

Alongside the result it returns a summary: file size, the sheets it found, row and column counts per sheet, and the number of distinct WRUA names, so an administrator can confirm the file is the one they meant.

Activity log

The Activity panel shows the most recent entries, newest first, each with who acted, what they did and when. Downloads, staged uploads and rejected uploads are all recorded.

The log is append-only JSON Lines in the runtime directory, which defaults to app/data/uploads/ and can be pointed elsewhere with RUNTIME_DIR. On a host with an ephemeral filesystem, Render among them, that directory does not survive a redeploy, so the log covers the current deployment and no further. Nothing else depends on it: the application works either way, it just forgets who downloaded what.

What is not built yet

Read this before promising anyone that data can be updated through the browser.

  • The page cannot upload. Every Replace button is rendered disabled, and the page shows the reason the inventory reports. It is shown rather than hidden so an administrator is not left wondering whether they are on the wrong screen.
  • The endpoint does not publish. POST /api/admin/upload/<source_id> validates a workbook and writes it into a staging directory. Its response says "published": false and states in words that the file is not live. Nothing reads the staged file afterwards.
  • No pipeline runs on the server. The step that turns a source workbook into the served CSV and GeoJSON files runs offline, and app/data_prep is excluded from the container image, so there is nothing on the server that could process an upload even if one were accepted.
  • Nothing triggers a reload. The application does have the machinery to pick up changed files: reload_data() re-reads everything and clears the GeoJSON cache, and a stamp file lets one worker tell the others to do the same. No code path calls it, because no code path publishes.

Until those are in place, replacing the data means preparing the files offline and deploying them, and this page is where you get the current copy and the format contract to work against.