Skip to content

Configuration

duckless init prints the first four. Keep them in a .envrc with direnv, or export them in your shell.

Variable Required
DUCKLESS_PROJECT yes project where jobs run
DUCKLESS_REGION no region of the jobs, default europe-west1
DUCKLESS_BUCKET yes work bucket
DUCKLESS_SA yes service account of the job VMs
DUCKLESS_IMAGE yes runner image used by run
DUCKLESS_NETWORK no network of the job VMs, default default
DUCKLESS_SUBNETWORK no subnetwork, default default; a full path for Shared VPC
DUCKLESS_EXTERNAL_IP no true to give job VMs an external IP (not needed with Private Google Access)
DUCKLESS_DUCKLAKE_INSTANCE no DuckLake catalog (Cloud SQL connection name); jobs attach it as lake
DUCKLESS_DUCKLAKE_DATA_PATH with the catalog where DuckLake writes table files, gcss://<bucket>/lake/

Set for every job, readable from SQL (${VAR}) and Python (os.environ):

Variable
DUCKLESS_JOB_ID the job id
DUCKLESS_BUCKET the work bucket
DUCKLESS_METRICS_URI where the runner writes its metrics
DUCKLESS_DUCKLAKE_INSTANCE, _DATA_PATH, _USER the catalog to attach, when the installation has one
GOOGLE_CLOUD_PROJECT the project

Settings you can pass with --env:

Variable Default
DUCKLESS_MEMORY_FRACTION 0.8 share of the VM memory given to DuckDB
DUCKLESS_GCS_GRPC true (false on Cloud Run Jobs) GCS over gRPC; false for HTTP
DUCKLESS_THREADS the VM’s vCPUs (the job’s vCPUs on Cloud Run Jobs) DuckDB threads
DUCKLESS_SCRATCH_DIR /mnt/disks/scratch where local SSD is mounted (spill goes under it)