Quick start
Before you start
Section titled “Before you start”- A Google Cloud project with billing enabled, and permission to enable APIs, create
service accounts and grant IAM roles on it (
roles/owneris the simple case). - Application Default Credentials:
gcloud auth application-default login. - Python 3.13 and uv (or pip).
- A subnetwork with Private Google Access in the region you use: job VMs have no
external IP. The
defaultnetwork has it on in most regions; check withgcloud compute networks subnets describe default --region <region> --format="value(privateIpGoogleAccess)".
1. Install
Section titled “1. Install”uv tool install duckless # or: pip install ducklessduckless --help2. Set up the project
Section titled “2. Set up the project”duckless init --project my-project --region europe-west1init enables the APIs it needs, then lets Infrastructure Manager create a work bucket, a
service account for the job VMs and a proxy for the runner image (details in
Infrastructure). It takes a few minutes and ends with
lines like these:
export DUCKLESS_PROJECT=my-projectexport DUCKLESS_REGION=europe-west1export DUCKLESS_BUCKET=my-project-duckless-workexport DUCKLESS_SA=duckless-runner@my-project.iam.gserviceaccount.comexport DUCKLESS_IMAGE=europe-west1-docker.pkg.dev/my-project/duckless-runner/tosun-si/duckless-runner:0.3.1Put them in your .envrc (or export them in your shell). Every other command reads them.
3. Check a machine
Section titled “3. Check a machine”duckless preflight --machine n2-highmem-16 --spotpreflight checks that the machine accepts the local SSD count DuckLess will attach, and
that the region has enough quota for it.
4. Run a first job
Section titled “4. Run a first job”Save this as hello.sql. It generates a small TPC-H dataset, writes it to your work
bucket as Parquet, and reads it back:
CALL dbgen(sf = 1);COPY lineitem TO 'gs://${DUCKLESS_BUCKET}/hello/lineitem' (FORMAT parquet, PER_THREAD_OUTPUT, FILE_SIZE_BYTES '256MB');SELECT l_returnflag, count(*) AS lines, round(sum(l_extendedprice), 2) AS revenueFROM read_parquet('gs://${DUCKLESS_BUCKET}/hello/lineitem/*.parquet')GROUP BY ALL ORDER BY ALL;duckless run hello.sql --machine n2-highmem-16 --spotThe CLI follows the job until it ends: about a minute for the VM to start, a few seconds to run. Then:
duckless result <job-id> # rows of the last SELECTduckless logs <job-id> # runner logs from Cloud LoggingClean up
Section titled “Clean up”duckless destroy --project my-project --force--force also deletes the work bucket if it still holds files.