# EC2 Provider

`--launcher EC2` provisions AWS instances. The control plane owns the
AWS SDK client and the decrypted credentials — the CLI never talks to
EC2 directly. The default `dispatch` is `daemon`: the instance's userdata
installs a `nymph` daemon that long-polls for work.

| Phase | What happens |
| ----- | ------------ |
| Provision | `RunInstances` via `@aws-sdk/client-ec2`, with your script base64-encoded into `UserData`. |
| Launch | Whatever the userdata script does — normally install and start the daemon. |
| Default dispatch | `daemon` |

## Where the launch happens

There are **two** launches here, and only the first one touches credentials.

| | Who acts | Credentials |
| --- | --- | --- |
| 1. Provision the box | control plane calls `RunInstances` | AWS keys, held by the control plane |
| 2. Get work to it | the nymph long-polls outbound | none — the daemon's own bearer token |

Step 2 is the same on every launcher. Step 1 is where EC2 differs from
SLURM and Kube: on those two a nymph can do the submitting itself
(`runner/slurm.rs`, `runner/kube.rs`), so the credential stays inside the
cluster. **There is no `ec2` runner in nymph** — `runner/mod.rs` registers
`process`, `docker`, `gvisor`, `slurm`, `kube`, `exec`, and `mock`, and
nothing else. So provisioning an instance is something only the control
plane can do.

> **Warning:** `aws_credentials` resolves from the Secret store into an `{ accessKeyId,
> secretAccessKey }` pair inside the control-plane process. The CLI genuinely
> never sees them — but "the CLI never sees them" is a weaker claim than
> "they never leave your account," and only the first one is true today.
> Scope the IAM user narrowly; see [Setting up a dedicated IAM
> user](#setting-up-a-dedicated-iam-user).

## Kwargs

These are the keys the EC2 launcher reads out of the merged
provider + `RunConfig` kwargs. Everything else round-trips on the row but
does not reach the SDK call.

| Field | Where it usually lives | Meaning |
| ----- | ---------------------- | ------- |
| `region` | provider | AWS region for the SDK client. |
| `aws_credentials` | provider | `{ accessKeyId, secretAccessKey }`. Omit to fall back to the SDK's default credential chain on the server. |
| `key_name` | provider | EC2 key-pair name → `KeyName`. |
| `security_group` | provider | One security group id → `SecurityGroupIds`. Use `security_group_ids` for a list. |
| `subnet_id` | provider | → `SubnetId`. |
| `iam_instance_profile_arn` | provider | → `IamInstanceProfile.Arn`. |
| `image_id` (or `ami`, or `ImageId`) | mode | **Required.** The launch throws without it. |
| `instance_type` (or `InstanceType`) | mode | Defaults to `t3.micro`. |

> **Warning:** The launcher builds a fixed `RunInstances` parameter set. There is no
> `spot_price`, no `availability_zone`, and no `dry` handling — setting
> those kwargs stores them on the row and changes nothing about the call.

Every instance is tagged `lakeshore=true`, `namespace=<ns>`, and
`provider=<name>`, plus any `tags` you pass on the launch body. User tags
that collide with those three reserved keys are dropped rather than
duplicated, which EC2 would reject.

## Seeding from an AWS profile

If the credentials already live in `~/.aws/credentials` +
`~/.aws/config`, skip the typing:

```bash
lakeshore aws discover                # what the CLI can see locally
lakeshore providers add aws-dev --from-aws-profile sandbox
```

`--from-aws-profile <name>` seeds `aws_credentials.accessKeyId`,
`aws_credentials.secretAccessKey`, `aws_credentials.sessionToken` (when
present), and `region` (when the profile or its config section sets one),
and infers `--launcher EC2`. Layer `--kwarg` on top for the rest:

```bash
lakeshore providers add aws-dev \
  --from-aws-profile sandbox \
  --kwarg key_name=lakeshore-ops \
  --kwarg security_group=sg-0a1b2c3d4e5f \
  --kwarg iam_instance_profile_arn='arn:aws:iam::…:instance-profile/lakeshore-worker'
```

**SSO profiles are rejected.** Any `sso_*` key in either file flags the
profile, and the launcher needs static keys. Run
`aws configure --profile <name>`, or pass the pair explicitly with
`--kwarg aws_credentials.accessKeyId=… --kwarg aws_credentials.secretAccessKey=…`.

Tab completion on `--from-aws-profile ` lists the local profile
names — see [Completion](https://lakeshore.dreamlake.ai/cli/completion).

> **Warning:** `--from-aws-profile` writes the keys straight into the provider kwargs.
> Fine for a scratch provider; move them into a Secret before the provider
> is shared.
>
> ```bash
> # The aws_keypair plaintext format is exactly `<access_key_id>:<secret_access_key>`.
> printf '%s:%s' AKIA… wJal… | lakeshore secrets add aws-dev-keys --kind aws_keypair
> lakeshore providers update aws-dev --kwarg aws_credentials='{"$secret":"aws-dev-keys"}'
> ```
>
> The launcher recognises `aws_keypair` secrets specially: it splits the
> plaintext on the first `:` and substitutes
> `{ accessKeyId, secretAccessKey }`. See
> [Secrets](https://lakeshore.dreamlake.ai/api/auth-and-secrets#secrets).

## Example

**.dreamrc tab:** The same shape in a local `.dreamrc`:

**CLI**

```bash
lakeshore providers add aws-us-east --launcher EC2 \
  --kwarg region=us-east-1 \
  --kwarg key_name=lakeshore-ops \
  --kwarg security_group=sg-0a1b2c3d4e5f \
  --kwarg iam_instance_profile_arn='arn:aws:iam::123456789012:instance-profile/lakeshore-worker'

lakeshore modes add gpu \
  --field provider=aws-us-east \
  --field instance_type=g5.xlarge \
  --field image_id=ami-0abcdef1234567890 \
  --field runner=docker \
  --field image=ghcr.io/lakeshore-py/cuda12.4-pytorch:2.4 \
  --field resources.gpu=1
```

**.dreamrc**

```yaml file=".dreamrc"
providers:
  aws-us-east: !providers.EC2
    region: us-east-1
    key_name: lakeshore-ops
    security_group: sg-0a1b2c3d4e5f
    iam_instance_profile_arn: arn:aws:iam::123456789012:instance-profile/lakeshore-worker

modes:
  gpu:
    provider: aws-us-east
    instance_type: g5.xlarge
    image_id: ami-0abcdef1234567890
    runner: docker
    image: ghcr.io/lakeshore-py/cuda12.4-pytorch:2.4
    resources: { gpu: 1 }
```

## Server-side routes

| Verb | Path | Result |
| ---- | ---- | ------ |
| `POST` | `/v1/namespaces/:ns/providers/:name/launch` | 202 `{ instanceId, state, providerName, launchedAt }` |
| `GET` | `/v1/namespaces/:ns/providers/:name/instances` | Instances tagged `lakeshore=true` for this namespace + provider |
| `DELETE` | `/v1/namespaces/:ns/providers/:name/instances/:id` | 204 |

`launch` requires a non-empty `script` (a complete bash userdata blob,
capped at 256 KiB) and takes optional `runConfigKwargs` and `tags`. The
launcher base64-encodes the script and hands it to `RunInstances`
verbatim — it never templates or interprets it.

`instanceId` is the EC2 instance id (`i-0abc1234`). `state` is normalized
from the SDK's `State.Name`: `pending → starting`, `running → running`,
`shutting-down`/`stopping → stopping`, `stopped`/`terminated`/`terminating
→ stopped`, anything else `unknown`. At launch time it is almost always
`starting` — poll `instances` to watch it become `running`.

AWS errors on a known list (`UnauthorizedOperation`, `AccessDenied`,
`InvalidAMIID.NotFound`, `InsufficientInstanceCapacity`, and friends)
come back as 502 with `{ error, awsCode, message }`; anything else is a
500, on the theory that it is our bug.

## Launching a daemon

For the `daemon` dispatch path you normally do not call `/launch`
yourself. `POST /v1/namespaces/:ns/daemons/launch` — the route behind
`lakeshore daemon launch` — renders the nymph bootstrap script (download
the binary, drop the unit, start it, point it at the control plane) and
passes it to this same launcher as the userdata:

```bash
lakeshore daemon launch --provider aws-us-east --name gpu --count 4 \
  --queue training --runner docker --keep-alive 1800 --wait
```

The bootstrap script's control-plane URL comes from
`LAKESHORE_PUBLIC_URL` on the server (falling back to
`LAKESHORE_SERVER_URL`, then `LAKESHORE_URL`, then localhost), so set
that on a deployed control plane or the daemons will try to reach
`http://localhost:8080`.

## Smoke test

```bash
lakeshore providers test aws-us-east
lakeshore providers instances aws-us-east
```

`providers test` posts a launch through the control plane. Streaming
stdout from the instance back to the CLI is not implemented — check the
instance over SSH, via `lakeshore daemon launch-log <id>` (EC2 serial
console), or by polling `instances`.

## Setting up a dedicated IAM user

For a fresh AWS account, the one-time path is: sign in as root → IAM →
create a programmatic user → attach a policy → create an access key →
wire it locally → register with Lakeshore.

1. Sign in as root at `https://signin.aws.amazon.com/` (with MFA).
2. **IAM → Users → Create user.** Leave console access unchecked —
   programmatic only.
3. **Permissions → Attach policies directly.** `AdministratorAccess` is
   the blunt option and covers everything the launchers touch. The
   minimum for EC2 alone is `ec2:RunInstances`, `ec2:Describe*`,
   `ec2:CreateTags`, `ec2:TerminateInstances`, plus `iam:PassRole` if you
   attach an instance profile.
4. **Security credentials → Create access key → Command Line Interface.**
5. On the success page, wire it up before navigating away — the secret is
   shown exactly once.

```bash
aws configure --profile lakeshore-dev
aws sts get-caller-identity --profile lakeshore-dev    # should print the IAM ARN

lakeshore providers add aws-us-east --from-aws-profile lakeshore-dev

# Then move the inline credentials into a Secret (see the callout above).
```

### Rotation

Static IAM keys never expire on their own. Rotate roughly every 90 days:

1. Create a **second** access key for the same user — AWS allows two
   active keys precisely so you can roll without downtime.
2. `aws configure --profile lakeshore-dev` with the new pair.
3. `printf '%s:%s' <new-id> <new-secret> | lakeshore secrets rotate <secret-name>`.
   Plaintext comes from `--from-file <path>` or piped stdin; there is no
   `--from-file -` form, so pipe it.
4. Smoke test: `lakeshore providers test aws-us-east`.
5. **Only then** deactivate the old key in the console, wait a day for
   stragglers to surface, and delete it.

Deactivate-then-delete is the point: a straggler on a deactivated key
fails visibly and reversibly. Deleting outright makes it fail
permanently and anonymously.

## OIDC federation — not built

Static keys in the control-plane database do not scale to "run in the
customer's AWS account." The clean answer is OIDC federation, the model
GitHub Actions uses: the control plane becomes an identity provider, the
customer creates an IAM role trusting it with `sub`/`aud` pinned to a
namespace, and each launch calls `sts:AssumeRoleWithWebIdentity` for
short-lived credentials. No keys at rest, per-launch audit.

None of that exists yet. It would need a JWKS + discovery endpoint on
the control plane with signing-key rotation, a `aws_role_arn` kwarg
mutually exclusive with `aws_credentials`, and an `AssumeRoleWithWebIdentity`
call in the EC2 launcher instead of the current credential passthrough.

AWS IAM Identity Center is not an alternative: SSO refresh is
fundamentally interactive, and a server has no browser session to renew
from. OIDC sidesteps that by making the control plane the issuer rather
than a consumer.
