App → database (fetch fresh on connect)¶
The same workload identity from App → services is the DB principal here — but the token lives inside the PostgreSQL password field, not an Authorization header. Fetch a fresh token per new physical connection; never cache a connection string.
Verified 2026-08-08
- Azure PG Flexible Server Entra auth (
https://ossrdbms-aad.database.windows.net/.default): learn.microsoft.com/…/postgresql/flexible-server/how-to-configure-sign-in-azure-ad-authentication - Cloud SQL IAM auth (Python Connector,
enable_iam_auth=True): cloud.google.com/sql/docs/postgres/connect-connectors - RDS IAM auth (Python) —
rds.generate_db_auth_token(...): docs.aws.amazon.com/AmazonRDS/latest/UserGuide/UsingWithRDS.IAMDBAuth.Connecting.Python.html
Per-cloud mechanics¶
Azure Postgres Flexible Server accepts an Entra access token as the password value. Audience is fixed: https://ossrdbms-aad.database.windows.net/.default. Token lifetime is 5–60 minutes; always fetch a fresh one at connect time.
import os, psycopg
from azure.identity import DefaultAzureCredential
_credential = DefaultAzureCredential()
_AAD_SCOPE = "https://ossrdbms-aad.database.windows.net/.default"
def _mint_password() -> str:
# Fresh token on every physical connect; do NOT cache the returned string
# outside this function. The credential object caches internally per scope.
return _credential.get_token(_AAD_SCOPE).token
def connect() -> psycopg.Connection:
return psycopg.connect(
host = os.environ["PGHOST"], # <server>.postgres.database.azure.com
port = 5432,
user = os.environ["PGUSER"], # user@tenant.onmicrosoft.com or the MI display name
password = _mint_password(),
dbname = os.environ["PGDATABASE"],
sslmode = "require",
)
Then use a pool with pool_recycle capped below the token lifetime:
from psycopg_pool import ConnectionPool
pool = ConnectionPool(
conninfo="", # ignored — use the callback
min_size=1, max_size=10,
max_lifetime=55 * 60, # 55 min < 60-min token cap
configure=lambda conn: None,
kwargs={},
connection_class=psycopg.Connection,
open=False,
)
# ... hand pool.getconn() the connect() above via `configure`
The pattern in examples/trust/passwordless/azure_pg.py uses psycopg's connection_factory hook so each physical connect goes through _mint_password() — no stale token in a pooled connection.
Cloud SQL for PostgreSQL uses IAM authentication via the Cloud SQL Python Connector. The connector handles IAM token acquisition + refresh automatically; you never touch the token string.
import os
from google.cloud.sql.connector import Connector
_connector = Connector(refresh_strategy="LAZY")
def connect():
# enable_iam_auth=True: the connector uses ADC (workload SA) to fetch
# a short-lived IAM auth token per connection; refresh is transparent.
return _connector.connect(
instance_connection_name = os.environ["INSTANCE_CONNECTION_NAME"], # "project:region:instance"
driver = "pg8000",
user = os.environ["DB_USER"], # SA email, e.g. chiron-chat@myproj.iam
db = os.environ["DB_NAME"],
enable_iam_auth = True,
)
Use with SQLAlchemy:
from sqlalchemy import create_engine
engine = create_engine("postgresql+pg8000://", creator=connect, pool_pre_ping=True,
pool_recycle=1800) # 30 min recycle; connector's own refresh is separate
The connector is the load-bearing piece — do not hand-roll a token loop around it; the LAZY refresh strategy is documented as the right default.
RDS for PostgreSQL uses IAM database authentication — rds.generate_db_auth_token(DBHostname, Port, DBUsername, Region) mints a signed URL (~15-minute TTL) that Postgres accepts as the password.
import os, boto3, psycopg2
_rds = boto3.client("rds", region_name=os.environ["AWS_REGION"])
def _mint_password() -> str:
return _rds.generate_db_auth_token(
DBHostname = os.environ["PGHOST"],
Port = 5432,
DBUsername = os.environ["PGUSER"],
Region = os.environ["AWS_REGION"],
)
def connect():
return psycopg2.connect(
host = os.environ["PGHOST"],
port = 5432,
database = os.environ["PGDATABASE"],
user = os.environ["PGUSER"],
password = _mint_password(), # fresh on every connect
sslrootcert = os.environ["RDS_SSL_ROOT"], # RDS-provided cert bundle
sslmode = "verify-full",
)
Pool cap must sit below the 15-min token TTL:
engine = create_engine(
"postgresql+psycopg2://",
creator=connect,
pool_recycle=14 * 60, # 14 min < 15-min IAM token TTL
pool_pre_ping=True,
)
IRSA is the identity underneath: boto3.client("rds") picks up the pod's IRSA role automatically; generate_db_auth_token() sigv4-signs the URL as that role. If your pod isn't annotated for IRSA, the AWS SDK falls back to node-role — usually the wrong identity, so make sure IRSA is wired.
The three failure modes to design against¶
-
Cached connection strings. Never store
postgresql://user:<token>@host/dbanywhere — as an env var, in a config object, or in an ORM'surlfield. Store the ingredients (host, user, DB name, workload identity via the SDK) and regenerate the password per connect. -
Long-lived pooled connections. A connection authenticated once, that lives past its token's expiry, is a bug waiting for the DB (or a middlebox, or an admin
pg_terminate_backend) to close it — and then reconnect logic can't mint a new token because you cached the old one. Setpool_recycleper the table below. -
Instantiating the credential inside the request path.
DefaultAzureCredential()/boto3.client("rds")/Connector(...)do non-trivial provider-chain setup. Build them once at process start and reuse. Per-request instantiation destroys the SDK's internal token cache and turns every request into an IMDS hit.
Pool-recycle cap per provider¶
| Provider | Token TTL | Recommended pool_recycle |
|---|---|---|
| Azure Entra → PG | 5–60 min (docs' guidance) | 55 min — a physical connection is dropped and re-opened before its token would practically expire |
| Cloud SQL Connector | Managed by the connector | 30 min is safe; connector's own refresh is independent of pool recycle |
RDS IAM generate_db_auth_token() |
15 min hard cap | 14 min — one minute of headroom below the token TTL |
The rule is: your pool's connection lifetime ≤ the token TTL minus a small safety margin. Cross that boundary and you get long-tail reconnect failures that only reproduce after N hours of production traffic.
Verify (validate-only)¶
Live verification needs real cloud creds + a real DB. The examples each ship a if __name__ == "__main__": smoke test that requires only the env vars listed at the top of the file.
Where the arc goes next¶
- Later in the trust arc — cross-account / cross-tenant federation (Entra guest tenants, GCP Workload Identity Federation from AWS, AWS
sts:AssumeRoleacross accounts). - RBAC / RLS composed on top of the attributed-writes pattern + these passwordless connections.