# GCE Provider

`--launcher GCE` provisions Google Compute Engine VMs. The control plane
owns the `@google-cloud/compute` client and the decrypted credentials —
the CLI never calls Google APIs directly. The default `dispatch` is
`daemon`: the VM's startup script installs a `nymph` daemon that
long-polls for work.

| Phase | What happens |
| ----- | ------------ |
| Provision | `instances.insert`, with your script attached verbatim as the `startup-script` metadata key. |
| Launch | Whatever the startup script does — normally install and start the daemon. |
| Default dispatch | `daemon` |

## Where the launch happens

Structurally identical to [EC2](/lakeshore/providers/ec2.md#where-the-launch-happens):
two launches, and only the first touches credentials.

| | Who acts | Credentials |
| --- | --- | --- |
| 1. Provision the VM | control plane calls `instances.insert` | GCP service-account key, held by the control plane |
| 2. Get work to it | the nymph long-polls outbound | none — the daemon's own bearer token |

There is no `gce` runner in nymph, so — as with EC2 — provisioning is
something only the control plane can do. Once the VM is up, the daemon on
it behaves exactly like any other.

> **Warning:** `dreamlake provider create` ships templates for `ssh`, `slurm`, `docker`
> (adopt an existing host) and `kube`, `ec2` (terraform to provision new
> infrastructure). **GCE has neither.** Register the provider against a VM you
> made yourself until a template lands.

## Kwargs

Required for a launch to succeed at all:

| Field | Where it usually lives | Meaning |
| ----- | ---------------------- | ------- |
| `project_id` | provider | GCP project. |
| `zone` | provider | e.g. `us-central1-a`. Also used to qualify machine and accelerator types. |
| `instance_type` | mode | Machine type, e.g. `n1-standard-8`, `a2-highgpu-1g`. |
| `image_id` **or** (`image_project` + `image_family`) | mode | The boot image. `image_project` + `image_family` expand to `projects/<p>/global/images/family/<f>`. |

Optional:

| Field | Meaning |
| ----- | ------- |
| `gcp_sa_json` | Service-account JSON for the SDK client. Omit to use the server's Application Default Credentials. |
| `name` | Instance name. Defaults to an auto-generated `lakeshore-<8 hex>` slug. |
| `boot_size` | Boot disk size in GB (number or string). |
| `accelerator_type` + `accelerator_count` | Attaches guest accelerators. Both are required together; setting them also forces `onHostMaintenance: TERMINATE`, which GCE requires. |
| `preemptible` | Exactly `true` sets `scheduling.preemptible`. |
| `tags` | A **list of strings** → network tags (firewall targeting). Not to be confused with the launch body's `tags`, which become labels. |
| `labels` | An object of label key/values, merged under the reserved ones. |

Every VM comes up on the `default` network with an ephemeral external IP
— the same shape `gcloud` produces with no flags — and with
`autoDelete: true` on the boot disk.

Labels always include `lakeshore=true`, `namespace=<ns>`, and
`provider=<name>`; user labels colliding with those three are dropped.
The `instances` listing filters on exactly those labels.

## Seeding from a gcloud configuration

```bash
lakeshore gcp discover                          # what the CLI can see locally
lakeshore providers add gcp-dev --from-gcloud-config research
```

`--from-gcloud-config <name>` seeds `project_id` (from `core/project`),
`region` (from `compute/region`), `zone` (from `compute/zone`), and
`gcp_sa_json` from the Application Default Credentials JSON at
`~/.config/gcloud/application_default_credentials.json` when it exists.
It also infers `--launcher GCE`.

`gcp discover` uses `gcloud config configurations list` when the `gcloud`
CLI is installed and otherwise walks `~/.config/gcloud/` directly. Tab
completion on `--from-gcloud-config ` lists the configuration names
— see [Completion](https://lakeshore.dreamlake.ai/cli/completion).

> **Warning:** `gcp_sa_json` lands verbatim in the provider kwargs. Move it into a
> Secret before the provider is shared:
>
> ```bash
> lakeshore secrets add gcp-dev-sa --kind gcp_sa_json --from-file ~/.config/gcloud/application_default_credentials.json
> lakeshore providers update gcp-dev --kwarg gcp_sa_json='{"$secret":"gcp-dev-sa"}'
> ```
>
> With no ADC file present the seeding skips that field and warns. Run
> `gcloud auth application-default login` first, or reference a stored
> secret directly. Note that `GOOGLE_APPLICATION_CREDENTIALS` is **not**
> consulted by the seeding path — only the ADC file.

## Example

**.dreamrc tab:** The same shape in a local `.dreamrc`:

**CLI**

```bash
lakeshore providers add gcp-us-central --launcher GCE \
  --kwarg project_id=my-project \
  --kwarg zone=us-central1-a

lakeshore modes add a100 \
  --field provider=gcp-us-central \
  --field instance_type=a2-highgpu-1g \
  --field image_project=deeplearning-platform-release \
  --field image_family=pytorch-latest-gpu \
  --field accelerator_type=nvidia-tesla-a100 \
  --field accelerator_count=1 \
  --field preemptible=true \
  --field runner=docker \
  --field image=ghcr.io/lakeshore-py/cuda12.4-pytorch:2.4 \
  --field resources.gpu=1
```

**.dreamrc**

```yaml file=".dreamrc"
providers:
  gcp-us-central: !providers.GCE
    project_id: my-project
    zone: us-central1-a

modes:
  a100:
    provider: gcp-us-central
    instance_type: a2-highgpu-1g
    image_project: deeplearning-platform-release
    image_family: pytorch-latest-gpu
    accelerator_type: nvidia-tesla-a100
    accelerator_count: 1
    preemptible: true
    runner: docker
    image: ghcr.io/lakeshore-py/cuda12.4-pytorch:2.4
    resources: { gpu: 1 }
```

## Server-side routes

| Verb | Path | Result |
| ---- | ---- | ------ |
| `POST` | `/v1/namespaces/:ns/providers/:name/launch` | 202 `{ instanceId, state, providerName, launchedAt }` |
| `GET` | `/v1/namespaces/:ns/providers/:name/instances` | VMs labelled `lakeshore=true` for this namespace + provider |
| `DELETE` | `/v1/namespaces/:ns/providers/:name/instances/:id` | 204 |

`instanceId` is the instance **name** — the auto-generated
`lakeshore-<hex>` slug or the `name` you set in kwargs — not a numeric
id. `launch` always returns `state: "starting"`, because
`instances.insert` is an async operation and the SDK returns before the
VM has a real status. Poll `instances` to see it reach `running`.

The listing uses GCE filter syntax on the reserved labels and prefers the
aggregated (cross-zone) list, falling back to a single-zone list when
`zone` is set.

## Launching a daemon

As with EC2, the `daemon` dispatch path normally goes through
`POST /v1/namespaces/:ns/daemons/launch` (`lakeshore daemon launch`),
which renders the nymph bootstrap script and hands it to this launcher as
the startup script. Set `LAKESHORE_PUBLIC_URL` on the control plane so
the script points daemons at the right host.

## Smoke test

```bash
lakeshore providers test gcp-us-central
lakeshore providers instances gcp-us-central
gcloud compute instances list
```

Streaming stdout from the VM back to the CLI is not implemented.

## Setting up a dedicated service account

GCP's default org policies actively block service-account key creation,
so this is heavier than the AWS equivalent. Step 1 needs the console once
per organization; the rest is `gcloud`.

### 1. Claim Organization Administrator (once per org)

A fresh Workspace org has an *unclaimed* GCP organization — nobody holds
`roles/resourcemanager.organizationAdmin`, so nobody can change org
policies. As Workspace super admin:

1. Open `https://console.cloud.google.com/iam-admin/iam?organizationId=`
   (find the id with `gcloud organizations list`).
2. Confirm the resource selector at the top reads your domain, not a
   project.
3. **+ GRANT ACCESS**, principal = your user, grant both
   `Organization Administrator` and `Organization Policy Administrator`.

Google honours your super-admin status as the implicit grantor here even
though no IAM binding gives you the permission.

### 2. Link billing

```bash
PROJECT=my-project
gcloud billing accounts list                       # find one with OPEN = True
gcloud billing projects link $PROJECT --billing-account=<ACCOUNT_ID>
```

A new project has no billing link even when the org has billing
accounts, and compute APIs refuse to enable without one
(`UREQ_PROJECT_BILLING_NOT_FOUND`).

### 3. Enable APIs

```bash
gcloud services enable \
  compute.googleapis.com \
  iam.googleapis.com \
  iamcredentials.googleapis.com \
  orgpolicy.googleapis.com \
  --project=$PROJECT
```

`orgpolicy.googleapis.com` is what makes step 4 possible from the CLI.

### 4. Override the SA-key policy for this project only

```bash
cat > /tmp/sa-key-policy.yaml <<'EOF'
constraint: constraints/iam.disableServiceAccountKeyCreation
booleanPolicy:
  enforced: false
EOF

gcloud resource-manager org-policies set-policy /tmp/sa-key-policy.yaml \
  --project=$PROJECT
rm /tmp/sa-key-policy.yaml
```

Wait 20–60 seconds for propagation, or the next step still fails on the
constraint. Every other project under the org stays protected.

### 5. Create the service account and a key

```bash
SA_NAME=lakeshore-dev
SA_EMAIL=${SA_NAME}@${PROJECT}.iam.gserviceaccount.com
KEY_PATH=~/.config/gcloud/${SA_NAME}.json

gcloud iam service-accounts create $SA_NAME \
  --display-name="lakeshore CLI" --project=$PROJECT

# Retry after ~10s if this reports the account does not exist (IAM propagation).
gcloud projects add-iam-policy-binding $PROJECT \
  --member="serviceAccount:$SA_EMAIL" --role="roles/owner"

mkdir -p ~/.config/gcloud
gcloud iam service-accounts keys create $KEY_PATH --iam-account=$SA_EMAIL
chmod 600 $KEY_PATH
```

`roles/owner` is the blunt option. A narrower equivalent is
`roles/editor` + `roles/compute.admin` + `roles/storage.admin` +
`roles/secretmanager.admin` + `roles/iam.serviceAccountAdmin` +
`roles/compute.networkAdmin`.

### 6. Point ADC at it, set the zone

```bash
gcloud auth activate-service-account --key-file=$KEY_PATH

# The CLI's --from-gcloud-config reads the ADC file, not
# GOOGLE_APPLICATION_CREDENTIALS.
cp $KEY_PATH ~/.config/gcloud/application_default_credentials.json
chmod 600 ~/.config/gcloud/application_default_credentials.json

gcloud config set project $PROJECT
gcloud config set compute/region us-central1
gcloud config set compute/zone us-central1-a

gcloud auth list                # the SA should be active
gcloud compute instances list   # 200, possibly empty
```

### 7. Register

```bash
lakeshore providers add gcp-us-central --from-gcloud-config default

lakeshore secrets add gcp-research-sa --kind gcp_sa_json --from-file $KEY_PATH
lakeshore providers update gcp-us-central \
  --kwarg gcp_sa_json='{"$secret":"gcp-research-sa"}'
```

### Rotation

Service-account JSON keys never expire on their own. Rotate roughly every
90 days:

1. Console → the SA → **Keys → Add key → Create new key → JSON**. GCP
   allows up to ten keys per SA, so rolling is easy.
2. Replace the local JSON and the ADC copy.
3. `lakeshore secrets rotate gcp-research-sa --from-file <new.json>`.
4. Smoke test: `lakeshore providers test gcp-us-central`.
5. Only then delete the old key. GCP has no deactivate step — deletion is
   immediate on confirm — so wait a day first and watch the audit logs
   for stragglers.
