Skip to content

Deployment and operations

Operator reference for running SWAG-DSS. Every statement here was checked against the repository. Where a figure depends on the target environment and cannot be established from the source, it is called out as unverified rather than estimated.

What the application is

Property Value
Runtime Python 3.11 (Dockerfile: FROM python:3.11-slim)
Web framework Flask 3.0.0
WSGI server gunicorn 23.0.0
WSGI entrypoint app:app (module-level app in app/app.py)
Container port 8000 (ENV PORT=8000, EXPOSE 8000)
Database none
Persistent storage none required; optional, for the admin audit log
Background workers none
Scheduled jobs none

Container start command, from the Dockerfile CMD:

gunicorn --bind 0.0.0.0:${PORT:-8000} --workers 2 --threads 4 --timeout 120 app:app

Shell form, so a platform-injected $PORT overrides the default of 8000. The container runs as the unprivileged user appuser.

Network requirements

Every front-end library is loaded from a public CDN at request time. None are vendored into app/static/ and the container carries no copies. The browser makes these requests, not the server, so what matters is what the user's browser can reach and what the page's Content-Security-Policy allows.

Host Serves If unreachable
unpkg.com Leaflet core and CSS the map does not render
cdn.jsdelivr.net Chart.js, Leaflet plugins, Bootstrap Icons charts and icons do not render
cdn.tailwindcss.com Tailwind CSS the page loads unstyled
fonts.googleapis.com web fonts fallback system fonts are used
www.gstatic.com Firebase JS SDK Google sign-in is unavailable

The application sets no Content-Security-Policy of its own. If one is imposed at the edge, these hosts must be allow-listed, or the dashboard renders as an unstyled page with no map and no charts. The application still returns HTTP 200 in that state, so a health check on / passes while the interface is unusable. Do not rely on the health check to catch this.

If Firebase sign-in is used, the deployed hostname must also be added in the Firebase console under Authentication, Settings, Authorized domains, or sign-in fails on the deployed host while working locally.

The application server itself needs no outbound internet access at runtime: all data it serves is on disk in the image. Outbound access is required only during the Docker build, to install system and Python packages.

System dependencies

The geospatial stack (GeoPandas, Fiona, Shapely, pyproj) links against GDAL, GEOS and PROJ. The image installs these Debian packages:

libgdal-dev libgeos-dev libproj-dev gdal-bin

and sets CPLUS_INCLUDE_PATH and C_INCLUDE_PATH to /usr/include/gdal so the Python packages build against them. Any deployment target that does not provide these libraries will fail at import time, not at request time. This is the reason to prefer a Docker-based platform over a generic Python platform.

Environment variables

Variable Required Purpose Behaviour when unset
SECRET_KEY yes for production Signs the Flask session cookie A random key is generated per process and a warning is written to stderr
ADMIN_USERNAME yes for production (unless using Firebase) Break-glass login username. This login is granted the admin role Credential login is disabled
ADMIN_PASSWORD yes for production (unless using Firebase) Break-glass login password Credential login is disabled
ADMIN_EMAILS no Comma-separated allowlist of email addresses granted the admin role. Matched case-insensitively, and only against an address Firebase reports as verified Nobody is an admin by email. Admin can still come from a Firebase custom claim
FIREBASE_API_KEY and other FIREBASE_* no Firebase authentication Firebase login is not offered
PORT no Listen port 8000 in the container
RUNTIME_DIR no Where the admin audit log is written app/data/uploads/, which is gitignored and does not survive a redeploy on an ephemeral host
CORS_ORIGINS no Comma-separated allowlist of origins that may read /api/* cross-origin. The dashboard is same-origin and needs no entry A built-in allowlist of the two known deployment hostnames
SESSION_COOKIE_SECURE no Set to 0 to allow the session cookie over plain HTTP, for local testing The cookie is marked Secure
BYPASS_AUTH no Set to 1 to skip the login gate and every role check; local development only Auth is enforced
FLASK_DEBUG no Set to 1 to enable the debugger when running app.py directly. Never in production Debug is off

See Data management for how ADMIN_EMAILS and the Firebase custom claim interact, and what admin access grants.

No credential has a committed fallback value. The two failure modes this produces are both silent at boot and visible only in use:

  • SECRET_KEY unset. The app starts and serves traffic. Because the container runs 2 gunicorn workers, each with its own random key, a session cookie signed by one worker is rejected by the other. Users appear to be logged out at random, and every restart or redeploy invalidates all sessions. The stderr warning at boot is the only direct signal.
  • ADMIN_USERNAME or ADMIN_PASSWORD unset. admin_login_enabled() in app/auth.py returns false, the login page does not render the credential form, and the POST branch rejects every submission, including one that omits the fields entirely. On a deployment with Firebase configured, users still sign in through Firebase. On one without it, there is no way to log in at all. That is the intended safe state, not a bug.
  • ADMIN_EMAILS unset and no Firebase custom claim set. Everyone resolves to viewer, so /admin and every /api/admin/* route returns 403. The dashboard is unaffected.

Set SECRET_KEY and a working sign-in path for any deployment that real users will log into. Supply credentials through the platform's secret store rather than the image or a committed file.

What happens at boot

  1. config.py is imported. It calls load_dotenv(), reads the environment once, and emits the SECRET_KEY warning if that variable is missing.
  2. create_app() in app/app.py builds the Flask app, applies ProxyFix so the external HTTPS scheme is honoured behind a proxy, restricts cross-origin /api/* reads to CORS_ORIGINS, and sets a 50 MB request-body cap (MAX_CONTENT_LENGTH).
  3. If BYPASS_AUTH=1 is set, a warning is written to stderr saying that authentication and role checks are disabled.
  4. init_firebase() runs. It is a no-op when Firebase is not configured.
  5. init_db() runs. This is the significant step: data.py reads every CSV and GeoJSON under app/data/ into a module-level in-memory dictionary. It is idempotent, and it runs once per worker process.
  6. The auth, pages and API blueprints are registered, along with the baseline security response headers (X-Content-Type-Options, X-Frame-Options, Referrer-Policy, and HSTS over HTTPS) and a before_request hook that stats a version stamp file so a worker can pick up a data change made by another worker. No code path writes that stamp today, so the hook is one stat() per request and nothing more.
  7. A line beginning Data loaded: is printed, reporting WRUA, reallocation, band and simulation row counts. Its absence from the logs means the data load did not complete. The counts in parentheses on that line are a leftover from an earlier data layout and carry no meaning; the totals are the useful part.

Missing data files are not fatal. load_data() substitutes an empty DataFrame or GeoDataFrame for any file it cannot find, so an incomplete app/data/ produces a running app that serves empty payloads and 200 responses. If the dashboard renders but shows no data, check that app/data/ reached the image.

Statelessness

Serving the dashboard writes nothing to disk and holds no cross-request state beyond the in-memory data cache.

  • CSV export builds the file in an in-memory buffer (app/api/exports.py).
  • PDF export writes into an in-memory buffer (app/api/report.py).
  • The download-all zip is assembled in memory (app/api/admin.py).
  • Sessions are client-side Flask cookies signed with SECRET_KEY.
  • The GeoJSON response cache is per-process and rebuilt on demand.

The one exception is the admin audit log, an append-only JSON Lines file under RUNTIME_DIR. It is written when an administrator downloads data or uploads a file. Nothing else reads it except the admin page and the upload endpoint's download-first check, and the application runs normally when it is missing or unwritable, so it does not make the app stateful in any way that affects serving.

Consequences for deployment: instances are interchangeable, can be scaled horizontally, and can be replaced at any time. No volume, no shared filesystem and no database are required. The only cross-instance requirement is that every instance share the same SECRET_KEY, so that a session issued by one instance is accepted by the others. Point RUNTIME_DIR at a persistent disk if the audit log needs to survive a redeploy.

Resource expectations

Resource What is known
Data on disk app/data/ is approximately 85 MB, of which the 13 files the admin inventory lists as served are 83.3 MB
Largest single files intervention_recommendations.csv at 32.9 MB, elevation_bands.geojson at 16.1 MB, and the two 48-month simulation CSVs at about 11.5 MB each
Memory The full dataset is parsed into pandas and GeoPandas objects and held for the process lifetime, once per worker. The container runs 2 workers, so the data is held twice. In-memory size is larger than the on-disk size, by a factor that depends on pandas dtype inference and has not been measured.
Startup time Dominated by the data load, since every file is parsed before the first request is served. Not measured on target hardware.
CPU Requests slice and aggregate tables already in memory. No hydrological computation happens at request time.

The gunicorn --timeout 120 is generous relative to normal request work and is sized for the heavier export and report endpoints.

Provision against measurement on the target instance type rather than against the figures above. Memory is the resource to watch, because it scales with the worker count.

Health checks

GET / returns 200. It is handled by pages.landing in app/pages.py and has no @login_required decorator, so it does not need a session and is suitable for an unauthenticated load-balancer health check. /about and /help are public on the same terms.

The page routes that require a session are /dashboard and /comparison, which redirect to /login when there is none, and /admin, which additionally requires the admin role. None of the three is usable as a health check.

Note that the health check succeeding does not imply the data loaded, since a missing dataset is non-fatal. Pair it with a check of the Data loaded: boot log line.

Documentation site

The /documentation/ routes serve a MkDocs site built into app/site/. The Docker build runs mkdocs build and writes the output there. If app/site/ is absent the route returns a 404 JSON hint and the rest of the application is unaffected, so the documentation build is optional and can be removed from the Dockerfile without breaking the application.

Tests

Run from app/, not from the repository root:

pip install pytest
cd app
python -m pytest tests -q

273 tests: 188 pass and 85 are expected failures. The suite sets BYPASS_AUTH=1 for itself and needs no environment configuration. app/tests/ is excluded from the container image.

The 85 are listed in app/tests/known_failures.txt and marked xfail(strict=True) by conftest.py. They are real assertions against figures the current dataset no longer produces, quarantined so the suite can still gate a merge: an unlisted test that fails is a normal failure and is loud, and a listed test that starts passing fails the build until its line is deleted, so the list can only shrink.

The working directory matters. conftest.py matches the quarantine against pytest node ids, which are relative to where pytest was invoked. Running python -m pytest app/tests from the repository root produces ids beginning app/tests/, none of which match the file, and the run reports 85 failures rather than 85 expected failures. Nothing is wrong with the code when that happens; the command is wrong.