Skip to content

Configuration reference

Available configuration options which can be set as environment variables:

Variable Type Default Description
ConnectionStrings__PostgreSQL string "" Connection string to the PostgreSQL database. See https://www.npgsql.org/doc/connection-string-parameters.html for options. Whenever this is set, it's also used to persist the ASP.NET Core Data Protection key ring so auth cookies/antiforgery tokens minted by one replica remain valid on another - independent of Authorization__IsEnabled, since Blazor Server's own circuit handshake relies on antiforgery regardless of whether OIDC auth is on. Vfps.dll migrate applies this context's migrations alongside the main one, so a separate migration Job covers both; if it is left unset there, the app fails at startup with relation "data_protection_keys" does not exist.
ForceRunDatabaseMigrations bool false Run database migrations as part of the startup. Only recommended when a single replica of the application is used.
DataProtection__Certificates__0__Path string "" Path to the certificate encrypting the Data Protection key ring at rest. A PKCS#12 archive, or a PEM certificate when __KeyPath is also set. Empty (the default) stores the key ring in plaintext in the same database as the pseudonyms - see Encrypting the Data Protection key ring. Index 0 protects newly created keys; further indices (__1, ...) are accepted for decryption only, which is what makes rotation possible.
DataProtection__Certificates__0__KeyPath string "" Path to the PEM private key belonging to __Path. Leave unset when __Path is a PKCS#12 archive, which carries its own key.
DataProtection__Certificates__0__Password string "" Password for the PKCS#12 archive or the encrypted PEM key. Empty means the file is not password-protected, which is the norm for a certificate mounted from a Kubernetes Secret.
ShutdownTimeout TimeSpan "0.00:00:25" How long the host waits for in-flight requests, Blazor circuits and background work to finish after SIGTERM before tearing down. Kept just under Kubernetes' default terminationGracePeriodSeconds of 30s so draining actually completes; raise both together (and add a preStop hook) if your clients need a longer drain - see Running more than one replica.
Tracing__IsEnabled bool false Enable distributed tracing support.
Tracing__ServiceName string "vfps" Tracing service name.
Tracing__RootSampler string "AlwaysOnSampler" Tracing parent root sampler. One of AlwaysOnSampler, AlwaysOffSampler, TraceIdRatioBasedSampler
Tracing__SamplingProbability double 0.1 Sampling probability to use if Tracing__RootSampler is set to TraceIdRatioBasedSampler.
Tracing__Otlp__Endpoint string "" The OTLP gRPC Endpoint URL traces are exported to (e.g. a Jaeger v2 or other OTLP-compatible collector).
Pseudonymization__Caching__Namespaces__IsEnabled bool false Set to true to enable namespace caching.
Pseudonymization__Caching__Pseudonyms__IsEnabled bool false Set to true to keep each original value's pseudonym in memory once it has been looked up, so repeated Create/Resolve calls for it skip the database. Only applies to namespaces that allow a single pseudonym per original value; CSV jobs are unaffected.
Pseudonymization__Caching__SizeLimit int 65534 Maximum number of entries in the cache. The cache is shared between the pseudonyms and namespaces.
Pseudonymization__Caching__AbsoluteExpiration D.HH:mm:nn 0.01:00:00 Time after which a cache entry expires.
Authorization__IsEnabled bool false Enable OIDC authentication and namespace-scoped authorization for the API and admin UI. Off by default: with this disabled, the API and admin UI are reachable without authentication and every namespace is fully accessible.
Authorization__Authority string "" The OIDC issuer/authority URL (e.g. a Keycloak realm URL).
Authorization__Audience string "" Audience validated for bearer-token (gRPC/REST API) callers.
Authorization__ClientId string "" Confidential client id used by the admin UI's Authorization Code flow.
Authorization__ClientSecret string "" Confidential client secret used by the admin UI's Authorization Code flow.
Authorization__RoleClaimType string "roles" Which claim on the validated token carries the caller's roles/groups.
Authorization__UsePushedAuthorizationRequests bool true Use RFC 9126 Pushed Authorization Requests when the authority advertises support for them. Set to false against older IdPs (e.g. pre-Quarkus Keycloak, before ~v19) whose PAR endpoint incorrectly rejects redirect_uri, which otherwise breaks every login with invalid_request: Invalid parameter: redirect_uri.
Authorization__AdminRoles__0, __1, ... string - Roles granting full access: all namespaces, plus namespace create/delete, plus managing every access grant. This is the only access setting that stays in configuration - somebody has to be an admin before there is any UI to grant anything from.
Authorization__GrantCacheDuration TimeSpan "0.00:00:30" How long a replica may answer permission checks from its in-memory snapshot of the access grants before re-reading them. The replica an admin makes a change on applies it immediately; this bounds how long the other replicas can still honour a grant that was just edited or revoked. Set to "0" to read the grants on every check instead, at the cost of a database round trip per permission check - including one per pseudonym Create.
Authorization__AccessTokens__IsEnabled bool false Let vfps issue its own bearer credentials - self-service personal access tokens, and admin-managed service accounts - for clients that can't get a token from the identity provider. See Access Tokens. Off by default: a long-lived static credential is strictly more exposed than a short-lived IdP-issued one, so a deployment that requires every caller to come through the identity provider leaves this unset and the whole feature, UI included, stays invisible. Requires Authorization__IsEnabled.
Authorization__AccessTokens__DefaultLifetime TimeSpan "90.00:00:00" Lifetime pre-filled when creating a token.
Authorization__AccessTokens__MaximumLifetime TimeSpan "365.00:00:00" The longest lifetime a token may be created with. Set to "0" to lift the cap - tokens still always expire, the creator just picks any date.
Authorization__AccessTokens__ExpiryWarningPeriod TimeSpan "14.00:00:00" How long before its expiry a token is marked Expires soon in the admin UI. Capped per token at a quarter of its lifetime, so a short-lived token isn't flagged from the moment it's created. Set to "0" to turn the marker off. Alerting on service-account tokens is configured separately - see Access Tokens.
Authorization__AccessTokens__UsageFlushInterval TimeSpan "0.00:01:00" How often each replica writes back the "last used" timestamps it has accumulated. Authentication only records them in memory, so this decides how stale that column can be, not how much work a request does. Purely informational - nothing about authentication or authorization reads it.

Per-namespace access is not configuration: it lives in the database and is managed from the admin UI's Access Control page. Earlier releases configured it through an Authorization__NamespaceRules__* section, which no longer exists - see Migrating from Authorization__NamespaceRules.

With Authorization__IsEnabled set, the API accepts bearer tokens only - see Authenticating API calls. The settings of the VOPRF client are listed on its own page.

CSV jobs and object storage

Variable Type Default Description
S3__IsEnabled bool false Enable CSV pseudonymization jobs (admin UI upload/download + the Hangfire job runner). Off by default - also requires ConnectionStrings__PostgreSQL, since Hangfire reuses the same database for its job storage. Note that Hangfire itself (and its /hangfire dashboard) is enabled by any deployment with a PostgreSQL connection string, independent of this setting - it also schedules internal housekeeping such as the pseudonym-count metric recompute.
S3__ServiceUrl string "" S3-compatible endpoint URL, e.g. a local MinIO instance or a real AWS/S3-compatible endpoint.
S3__AccessKey string "" Access key for the bucket above.
S3__SecretKey string "" Secret key for the bucket above.
S3__Bucket string "" Bucket CSV job input/output files are stored in.
S3__Region string "us-east-1" Region passed to the S3 client.
S3__ForcePathStyle bool true Path-style addressing (https://host/bucket/key) rather than virtual-hosted-style - required for MinIO and most non-AWS S3-compatible stores.
S3__PresignedUrlExpiry TimeSpan "0.00:15:00" How long presigned upload/download URLs remain valid.
S3__ObjectRetentionDays int 30 Days before an S3 bucket lifecycle rule expires (deletes) a CSV job's input/output objects. Applied on startup, scoped to the csv-jobs/ prefix. Job records themselves are never deleted, so without this the original, unpseudonymized input would otherwise live in the bucket forever. Set to 0 to leave the bucket's lifecycle configuration untouched. Overwrites the bucket's entire existing lifecycle configuration - use a bucket dedicated to vfps.
S3__AllowedOrigins string[] [] Origins (e.g. https://vfps.example.org) allowed to PUT/GET objects directly against the bucket via presigned URLs, applied as an S3 bucket CORS rule on startup. The browser talks to the bucket on a different origin than vfps itself, so without this the browser blocks the upload with a CORS error before it reaches S3. Empty (the default) leaves the bucket's CORS configuration untouched. Overwrites the bucket's entire existing CORS configuration - use a bucket dedicated to vfps.
CsvProcessing__PseudonymizeBatchSize int 1000 How many rows' worth of values go into one batched upsert round trip when pseudonymizing a CSV job. Doesn't apply to de-pseudonymization, which resolves concurrently instead.
CsvProcessing__StalledJobThreshold TimeSpan "0.00:10:00" How long a CSV job can sit Running with no progress update before it's marked Stalled (its worker likely crashed, lost its database connection, or was killed by an app restart).
CsvProcessing__ProcessJobs bool true Whether this instance runs CSV jobs, as opposed to only accepting them. Enqueueing and the Hangfire dashboard work either way - this controls only which Hangfire queues the instance serves (false serves internal housekeeping such as the pseudonym-count recompute, but not the queue CSV jobs land in), which is what allows dedicated worker pods: set it to false on the pods serving the API and admin UI, and true on a second deployment of the same image. If every instance sets it to false, jobs queue forever with no error.
CsvProcessing__WorkerCount int 4 How many CSV jobs one replica processes concurrently. Pinned rather than left at Hangfire's own default (min(cores * 5, 20)), which knows nothing about the shared Npgsql connection pool: a de-pseudonymizing job resolves each chunk via up to 20 concurrent lookups, so 20 workers would be up to 400 concurrent connection requests from one replica against a pool whose default maximum is 100. Raise it in step with Maximum Pool Size, and remember it multiplies by replica count against one shared database.
CsvProcessing__JobServerShutdownTimeout TimeSpan "0.00:00:15" How long the Hangfire job server waits for in-flight jobs to wind down on shutdown. A CSV job can't checkpoint and resume, so this doesn't let one finish - it buys time to unwind cleanly and record its own outcome. Must stay comfortably below ShutdownTimeout.
CsvProcessing__OrphanedJobRecoveryDelay TimeSpan "0.00:02:00" How long a job orphaned by a replica disappearing (rolling upgrade, OOM kill, node failure) waits before another replica picks it up and reprocesses it from the start. Hangfire's own default is 5 minutes, which races uncomfortably closely with CsvProcessing__StalledJobThreshold; 2 minutes makes re-dispatch the reliable winner, so an upgrade costs a job a couple of minutes rather than a Stalled status. Hangfire's heartbeat and server-check intervals are derived from this; values below 30s are clamped.
CsvProcessing__MissingValuePlaceholders string[] ["NA", "NULL"] Source values (matched case-insensitively, after trimming) treated as "no value" and passed through to the output unchanged instead of being pseudonymized/de-pseudonymized. A blank/whitespace-only cell is always treated this way regardless of this setting. Set to [] to only skip genuinely blank cells.