Skip to content

Quick start

  • A Google Cloud project with billing enabled, and permission to enable APIs, create service accounts and grant IAM roles on it (roles/owner is the simple case).
  • Application Default Credentials: gcloud auth application-default login.
  • Python 3.13 and uv (or pip).
  • A subnetwork with Private Google Access in the region you use: job VMs have no external IP. The default network has it on in most regions; check with gcloud compute networks subnets describe default --region <region> --format="value(privateIpGoogleAccess)".
Terminal window
uv tool install duckless # or: pip install duckless
duckless --help
Terminal window
duckless init --project my-project --region europe-west1

init enables the APIs it needs, then lets Infrastructure Manager create a work bucket, a service account for the job VMs and a proxy for the runner image (details in Infrastructure). It takes a few minutes and ends with lines like these:

Terminal window
export DUCKLESS_PROJECT=my-project
export DUCKLESS_REGION=europe-west1
export DUCKLESS_BUCKET=my-project-duckless-work
export DUCKLESS_SA=duckless-runner@my-project.iam.gserviceaccount.com
export DUCKLESS_IMAGE=europe-west1-docker.pkg.dev/my-project/duckless-runner/tosun-si/duckless-runner:0.3.1

Put them in your .envrc (or export them in your shell). Every other command reads them.

Terminal window
duckless preflight --machine n2-highmem-16 --spot

preflight checks that the machine accepts the local SSD count DuckLess will attach, and that the region has enough quota for it.

Save this as hello.sql. It generates a small TPC-H dataset, writes it to your work bucket as Parquet, and reads it back:

CALL dbgen(sf = 1);
COPY lineitem TO 'gs://${DUCKLESS_BUCKET}/hello/lineitem'
(FORMAT parquet, PER_THREAD_OUTPUT, FILE_SIZE_BYTES '256MB');
SELECT l_returnflag, count(*) AS lines, round(sum(l_extendedprice), 2) AS revenue
FROM read_parquet('gs://${DUCKLESS_BUCKET}/hello/lineitem/*.parquet')
GROUP BY ALL ORDER BY ALL;
Terminal window
duckless run hello.sql --machine n2-highmem-16 --spot

The CLI follows the job until it ends: about a minute for the VM to start, a few seconds to run. Then:

Terminal window
duckless result <job-id> # rows of the last SELECT
duckless logs <job-id> # runner logs from Cloud Logging
Terminal window
duckless destroy --project my-project --force

--force also deletes the work bucket if it still holds files.