sky-lattice/PLAN.md
VG 611ad214fe Initial Sky Lattice scaffold: blueprint, platformctl, and docs.
Encode intent/decision/plan workflow for Snowflake platform delivery so engagements share a durable recipe instead of one-off LLM chats.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-15 01:48:36 -04:00

1035 lines
42 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
name: Snowflake Bootstrap Strategy
overview: Internal blueprint repo (modules, schemas, wizard, policies, skills) plus per-customer repos. Wizard drives greenfield/brownfield intent. Customer env gets repo+CI+secrets+Snowflake automation user; blueprint IP stays pinned on operator side by default.
todos:
- id: blueprint-repo-layout
content: Finalize blueprint repo contents (modules, schemas, wizard catalogs, policies, skills, customer-repo template)
status: in_progress
- id: wizard-catalogs
content: Design greenfield and brownfield wizard question catalogs and answer→intent/decision mapping
status: pending
- id: customer-landing
content: Define customer-env bring-up (Snowflake service user, secrets, TF state, git/CI) and default delivery mode A
status: pending
- id: cli-packaging
content: Package platformctl as Python (pipx/uv); publish private; CLI never depends on live access to operator blueprint working tree
status: pending
- id: lifecycle-edit
content: Define post-apply lifecycle — resumable wizard, intent edit→plan→apply as default; scoped destroy as exception
status: pending
- id: low-friction-ux
content: Design low-friction UX — 5 commands, Snowflake vocabulary, LLM skill front door, TF-native plans, recipes, doctor/explain
status: pending
- id: intent-model
content: Define intent/desired-config model above raw TF (blueprint params + customer deviations)
status: pending
- id: decision-store
content: Design decision/rationale store linked to intent paths and overrides
status: pending
- id: observe-reconcile
content: Design observe path — TF state + live Snowflake inventory + drift classification
status: pending
- id: deep-dive-next
content: Next deep-dive — wizard catalog fields vs intent schema (so questions map cleanly)
status: pending
isProject: false
---
# Snowflake Platform Tool: Layered Design (Revised)
## Corrections to prior framing
Two important clarifications from you:
1. **Session N+1 is not empty.** You pass Terraform state (and related IaC), so the LLM can reconstruct *what exists*. What does **not** travel with state is **why** a choice was made (rationale, rejected alternatives, temporary compromises, client constraints).
2. **“LLM-driven” ≠ “LLM owns the work.”** Execution and ongoing management are already (or should remain) **IaC / tools**. The LLM is how developers *drive* generation and evolution of that code today — with weak durable intent/decision memory between sessions.
Prior plan over-indexed on “empty context” and under-specified the real stack. This revision treats the problem as a **multi-layer system**, not a single bootstrapper slogan.
---
## Precise problem statement
| What you already have (or can pass in) | What decays between sessions |
|---|---|
| TF state / config (what is managed) | Decision rationale (why this shape) |
| Live Snowflake (what exists) | Intent vs exception (is this deviation permanent?) |
| LLM + developer loop (how changes get authored) | Stable operating procedures as *bound* skills over a schema |
| Best practices in peoples heads / docs | Machine-checkable policy + linked ADRs |
So the product is not “replace Terraform.” It is:
**an intent + decision + policy layer on top of IaC**, with LLM as a skilled driver that must read/write those layers — not only TF state.
---
## Layer stack (this is the real work)
```mermaid
flowchart TB
subgraph L7 [L7_OperatorUX]
chat[LLM_Chat_Skills]
cli[CLI_PR_Workflow]
end
subgraph L6 [L6_PolicyVerify]
policy[PolicyAsCode]
verify[PostApplyAssertions]
end
subgraph L5 [L5_PlannerCodegen]
planner[IntentDelta_to_TFPlan]
codegen[ModuleAware_Codegen]
end
subgraph L4 [L4_DecisionKnowledge]
adr[DecisionRecords]
skills[OpsSkills_BestPractices]
end
subgraph L3 [L3_IntentModel]
blueprint[Blueprint_vN]
intent[CustomerIntent]
overrides[Overrides]
end
subgraph L2 [L2_Observe]
tfstate[TerraformState]
inventory[LiveSnowflakeInventory]
drift[DriftClassifier]
end
subgraph L1 [L1_Execute]
tf[Terraform_Apply]
snow[SnowflakeAccount]
end
chat --> intent
chat --> adr
chat --> planner
cli --> planner
skills --> chat
blueprint --> intent
overrides --> intent
intent --> planner
adr --> planner
tfstate --> drift
inventory --> drift
drift --> planner
planner --> codegen --> tf --> snow
tf --> tfstate
snow --> inventory
policy --> planner
verify --> snow
```
Each layer has its own schema, APIs, and failure modes. They should be designed deliberately — not collapsed into “a bootstrap tool.”
---
### L1 — Execution (already largely exists)
**Owns:** applying changes to Snowflake
**Artifacts:** Terraform modules/resources, providers, backends, CI apply
**Non-goals:** remembering why; encoding client intent in a portable form
Assumption going forward: **Terraform remains the executor** (you already pass TF state into sessions). Other engines only matter if you later choose them; do not reopen that as the center of the product.
---
### L2 — Observe / reconcile
**Owns:** truthful picture of managed vs live vs drifted
Inputs:
- Terraform state + config
- Live account inventory (SHOW / INFORMATION_SCHEMA / account usage as needed)
Outputs:
- Resource inventory joined to TF addresses
- Drift classes: `in_sync`, `drifted_managed`, `unmanaged_live`, `in_state_missing_live`
- Adoption candidates for brownfield
This layer answers “where things are.” You already approximate it by feeding TF state to the LLM; inventoriing live Snowflake closes blind spots TF does not see (manual grants, console clicks, out-of-band objects).
---
### L3 — Intent model (above raw TF)
**Owns:** *what we mean to run for this customer*, parameterized
Artifacts (illustrative):
- `blueprint.lock` — pinned platform blueprint version
- `intent.yaml` — domains, envs, data products, role assignments, warehouse profiles
- `overrides.yaml` — explicit deviations from blueprint defaults
Raw TF is a **compiled artifact** of intent. Today the LLM often edits TF directly; that works, but intent stays implicit inside HCL. Making intent explicit is what makes “tweak per client” repeatable without re-deriving the whole graph from chat.
Deviations live here as data, not as undocumented HCL quirks.
---
### L4 — Decision / knowledge (the real N+1 gap)
**Owns:** *why* — durable across sessions even when TF state is provided
Artifacts:
- Decision records linked to resources or intent keys, e.g.:
- `decision_id`, `applies_to` (TF address or intent path), `status` (active/superseded), `rationale`, `alternatives_rejected`, `expires_on`, `client_constraint`
- Curated **ops skills / best-practice packs** (versioned markdown/YAML the LLM must load)
- Mapping: blueprint rule → allowed override types → required decision fields
Example of what TF state cannot tell the next session:
- “We kept a flat role for finance because their IdP groups land weekly; revisit after SCIM.”
- “Prod shares DEV database naming temporarily for migration; remove after cutover date.”
- “Blueprint says account-per-env; this client insisted on database-per-env for cost.”
Without L4, every N+1 session can rebuild *structure* from state and still **re-litigate or contradict** prior judgment.
---
### L5 — Planner / codegen
**Owns:** turning intent deltas (+ decisions) into concrete IaC changes
Flow:
1. Load blueprint + intent + overrides + decisions
2. Load observe/drift report
3. Compute target module graph / resource set
4. Emit TF changes (or a plan preview) — **schema-validated**
5. Attach/update decision records for any override touched
LLM role here: assist codegen and explain diffs **inside this pipeline**, not free-form “write some Snowflake SQL.” Guardrails: module allowlist, forbidden patterns, required decision when overriding policy.
---
### L6 — Policy / verify
**Owns:** machine-checkable best practices (not only prose skills)
Examples:
- No `ACCOUNTADMIN` for service users
- Raw zone writable only by loader roles
- Env isolation invariants
- Naming / tagging requirements
- Override must reference an active decision record
Runs on: pre-plan validate, PR checks, post-apply assertions.
Skills (L4) teach the LLM; policy (L6) **fails the build** when the LLM or a human skips the rules.
---
### L7 — Operator UX
**Owns:** how developers drive the system
- CLI: `discover`, `intent validate`, `plan`, `apply`, `drift`, `decision add`
- LLM skills: load L3+L4+L2 context automatically; propose intent/override/decision patches; call planner
- PR workflow: intent/decision/TF in one reviewable unit
This is where “LLM-driven” correctly lives — as the **driver**, with IaC as the **engine**.
---
## What “repeatability with deviation” means in this stack
```text
Blueprint_vN
+ Customer intent (params)
+ Overrides (structured deviations)
+ Decisions (why those overrides exist)
→ Planner → Terraform → Snowflake
→ Observe → Drift → next change
```
Same steps, different inputs. Not the same HCL pasted every time; not a blank-slate LLM redesign every time either.
---
## Implementation breadth (why this feels “many layers”)
Rough build slices (can be separate workstreams):
1. **Customer project layout** — conventions for TF root, intent, decisions, skills pin
2. **Blueprint modules** — your recommended infra as versioned TF modules (not one-off roots)
3. **Intent schema + compiler** — intent → module inputs (deterministic first)
4. **Decision store format + linking** — resource/intent annotations
5. **Discover/drift** — state + live inventory join
6. **Policy pack** — OPA/Conftest/custom validators on intent+TF
7. **LLM skill pack** — Cursor/agent skills that require reading L2L4 before editing
8. **Brownfield adopt** — import live → intent proposal → unmanaged list
9. **Verification suite** — Snowflake-side integration checks per customer
10. **Multi-customer packaging** — how blueprints version and how clients pin upgrades
This is product/platform engineering, not a weekend script. The earlier “bootstrapper” framing was the tip of that iceberg.
---
## Recommended discussion order (next deep-dives)
Do not design all layers at equal depth at once. Suggested sequence:
1. **L3 Intent model** — what must be expressible without dropping to raw TF
2. **L4 Decision store** — schema that captures the why you lose today
3. **L5 Planner contract** — how intent patches become module-safe TF diffs
4. **L2 Observe** — drift model given you already use TF state
5. **L6 Policy** — which best practices become hard gates first
6. **L7 Skills** — only after schemas exist so skills have something binding to edit
---
## Verdict (revised)
- You are already closer than “LLM owns infra”: **IaC owns execution; LLM drives authoring.**
- Passing TF state into session N+1 solves **topology memory**, not **decision memory**.
- The useful internal product is a **layered intent/decision/policy system over Terraform**, with discover/drift and LLM skills as interfaces — not a single high-level “bootstrap tool” narrative.
- Next value is to **deep-dive one layer at a time** and nail schemas/contracts before writing a lot of code.
---
## Blueprint repo — what lives there
The blueprint repo is **your IP / product core**. Customer repos consume a pinned version of it; they do not fork a copy of everything.
```text
snowflake-platform-blueprint/
modules/ # Terraform modules (executor building blocks)
environment/ # env isolation pattern
database_zones/ # landing/raw/curated/marts (or your real zones)
rbac/ # role hierarchy + grant patterns
warehouses/
service_principals/
monitoring/ # resource monitors, basic alerts stubs
schemas/
intent.schema.json # what intent.yaml must look like
decision.schema.json
discover-report.schema.json
defaults/
intent.greenfield.yaml # starter intent
intent.brownfield.yaml
warehouse_profiles.yaml
access_profiles.yaml # e.g. standard vs restricted
wizard/
catalog.greenfield.yaml # ordered questions, branching, defaults
catalog.brownfield.yaml # starts with discover-driven questions
mapping.yaml # answer_id → intent path / override / decision prompt
policies/ # hard gates (Conftest/OPA/custom)
no_accountadmin_services.rego
env_isolation.rego
override_requires_decision.rego
skills/ # LLM operator packs (Cursor/agent skills)
platform-ops/SKILL.md
rbac-best-practices.md
brownfield-adopt.md
templates/
customer-repo/ # skeleton copied by `init`
intent.yaml
decisions/.gitkeep
terraform/
.github/workflows/ # optional CI template for customer
tests/
reference-customer/ # golden intent → expected plan fixtures
CHANGELOG.md # blueprint v1, v2, …
VERSION
```
**Rule of thumb**
| In blueprint repo | In customer repo |
|---|---|
| Modules, schemas, wizard catalogs, policies, skills, defaults | Filled `intent.yaml`, `decisions/`, TF root that *calls* modules, state backend config, connection secrets refs |
| Versioned releases (`v3.2.0`) | `blueprint: v3.2.0` pin |
| Your best-practice brain | This clients answers and exceptions |
---
## Wizard — yes, first-class
The wizard is the primary greenfield/brownfield onboarding UX. It is **not** free-form chat inventing infra; it is a **question catalog** over the blueprint that writes `intent.yaml` + prompts for `decisions/` when answers deviate from defaults.
### Shape
```mermaid
flowchart LR
mode{Greenfield_or_Brownfield}
discover[DiscoverLiveAccount]
ask[WizardQuestionCatalog]
intent[Write_intent.yaml]
decide[PromptDecisionIfOverride]
plan[platformctl_plan]
apply[ApplyInAllowedContext]
mode -->|greenfield| ask
mode -->|brownfield| discover --> ask
ask --> intent --> decide --> plan --> apply
```
### Greenfield wizard (examples)
Ordered questions from `wizard/catalog.greenfield.yaml`:
1. Customer name / account identifiers
2. Env model: account-per-env vs db-per-env (default from blueprint)
3. Domains / data products (repeatable)
4. Access profile per domain (standard / restricted)
5. Warehouse profile
6. Identity: SSO now / later / service users only
7. Anything that contradicts a default → **forced decision capture** (rationale + optional expiry)
Answers compile via `wizard/mapping.yaml``intent.yaml`.
CLI shape:
```bash
platformctl wizard --customer acme --mode greenfield
# or non-interactive from a saved answers file:
platformctl wizard --answers acme-answers.yaml
```
LLM can *run* the same catalog conversationally (“ask next unanswered question”), but the catalog is the contract — same outputs whether human CLI or chat.
### Brownfield wizard
1. Connect + `discover`
2. Show match % / unmanaged / conflicts
3. Questions become **adoption choices**: manage / ignore / migrate-later per object group
4. Propose intent draft from what matched
5. Require decisions for keep-legacy and naming conflicts
6. `adopt` → import selected into TF → plan residual gaps
Same wizard engine; different catalog + discover as step 0.
---
## Bringing this into a new customer environment
Split **operator toolkit** (yours) from **runtime footprint** (what must exist in the customers world).
### Default delivery model (recommended)
**You keep the blueprint + wizard.** The customer receives a **customer project** that can plan/apply against *their* Snowflake, usually from *their* Git + CI — without needing your full blueprint source if modules are consumed as a versioned package (git tag, private module registry, or vendored lock).
```mermaid
flowchart TB
subgraph your_side [YourSide_Operator]
blueprint[BlueprintRepo]
wizard[Wizard_CLI_or_Chat]
end
subgraph customer_side [CustomerEnvironment]
crepo[CustomerGitRepo]
ci[CustomerCI]
tfstate[TFStateBackend]
sf[SnowflakeAccount]
end
wizard -->|generates_or_updates| crepo
blueprint -->|pinned_modules| crepo
crepo --> ci --> tfstate
ci -->|terraform_apply| sf
```
**What must be set up once in the customer environment**
1. **Snowflake bootstrap identity** — service user + key pair (or OAuth), roles enough for Terraform (often a dedicated automation role; initial ACCOUNTADMIN bootstrap is one-time manual)
2. **Secrets** — private key / cloud creds in customer secret store (GitHub Actions secrets, Vault, etc.) — never in git
3. **TF state backend** — customer-controlled (S3/Azure/GCS + lock), or your orgs backend if *you* operate the account under contract
4. **Customer git repo** — the generated project from `init` / wizard
5. **CI workflow** — plan on PR, apply on merge/approval (template from blueprint `templates/customer-repo`)
**What does *not* need to live in the customer env**
- Wizard question catalogs (can run on your laptop during engagement)
- Full policy authoring / skill authoring
- Blueprint module *source* (if you publish versioned modules they consume read-only)
### How day-1 “bring-up” works in practice
1. Customer creates (or you create) empty Snowflake account / grants you access
2. One-time manual: create Terraform service user + store key in their secrets
3. You run wizard (locally or internal portal) → produces customer repo contents
4. Push customer repo to **their** GitHub/GitLab (or your managed repo if contract says so)
5. Wire secrets + state backend
6. CI runs `plan`/`apply` into their account
7. Ongoing changes: PR to customer repo (intent/decision/TF); wizard/LLM used by your team as operators
### Alternate operating modes (when needed)
| Mode | When | Footprint in customer |
|---|---|---|
| **A. Customer-hosted CI (default)** | Normal consulting delivery, customer owns account | Repo + secrets + state + Snowflake service user |
| **B. You operate apply** | You manage platform under retainer | Same Snowflake identity; state/CI may stay in your org |
| **C. Fully offline customer** | No external module fetch | Vendor/pin blueprint modules into customer repo at release time |
Avoid shipping the entire blueprint monorepo into every customer unless they require air-gapped ownership — that blurs IP and versioning.
### “Tool in customer environment” checklist (handoff)
- [ ] Snowflake automation user created
- [ ] Secrets installed in CI
- [ ] State backend ready
- [ ] Customer repo pushed with blueprint pin
- [ ] First successful `plan` in CI
- [ ] First successful `apply`
- [ ] Drift job scheduled or documented
- [ ] Runbook: how to add a domain / raise a decision / upgrade blueprint pin
---
## CLI access model and packaging
### Will the CLI reach into “my” repo from the customer environment?
**No — not if designed correctly.** That would be brittle (VPN, private git auth, IP leakage) and is the wrong dependency direction.
Correct dependency direction:
```text
Published artifacts (versioned)
├── platformctl (CLI package on private PyPI / uv tool)
└── blueprint modules + schemas + wizard catalogs (versioned release, e.g. v3.2.0)
Customer environment / CI only needs:
├── customer git repo (intent, decisions, TF root)
├── platformctl (optional; CI may only need terraform)
└── ability to fetch pinned blueprint vX (registry / git tag / vendored copy)
```
The CLI works against the **local customer project directory** (and Snowflake + TF state). It does **not** SSH into your laptop or clone your live blueprint working tree.
| Who runs what | Accesses |
|---|---|
| You on laptop (wizard during engagement) | Local checkout or cached blueprint package + writes customer repo |
| Customer CI `terraform plan/apply` | Customer repo + pinned TF modules only (often **no** platformctl) |
| Customer engineer with platformctl (optional) | Same: local customer repo + installed CLI + pinned blueprint package from registry |
If the customer is air-gapped: **vendor** blueprint `v3.2.0` into the customer repo at release time (mode C). Still no live link to your private monorepo.
### Packaging recommendation
**Default: Python package, installed via `pipx` or `uv tool install`.**
Why Python over npm/Go for this tool:
- Snowflake ecosystem is Python-native (`snowflake-connector-python`, many internal data-platform scripts)
- Excellent YAML/JSON schema, Jinja, Click/Typer CLIs
- Easy private publish (Azure Artifacts, AWS CodeArtifact, GitHub Packages, private PyPI, or `pip install git+https://…`)
- `pipx`/`uvx` gives an isolated CLI without polluting customer project venvs
| Option | Fit | Use when |
|---|---|---|
| **Python + pipx/uv (recommended)** | Best overall | Your team already does Snowflake/Python |
| **Go single binary** | Best “download one file” DX | Strict customer lockdown, no Python allowed on jump hosts |
| **npm** | Weak fit | Only if your org is JS-only; TF/Snowflake glue will feel awkward |
| **Snowflake CLI plugin** | Niche | If you standardize on `snow` CLI everywhere; still thin wrapper over same Python core |
| **Docker image** | Good CI companion | `ghcr.io/yourorg/platformctl:3.2.0` for customer pipelines without local install |
Suggested package split:
1. **`platformctl`** — CLI (wizard, validate, plan wrapper, decision helpers, discover)
2. **`platform-blueprint`** (or git-tagged module repo) — versioned content the CLI resolves by `blueprint: v3.2.0` in intent
Publish both to **your** private registry. Customer CI authenticates to that registry *or* you vendor the blueprint pin into their repo so they need zero access to your org.
### Minimal customer footprint (preferred)
Many engagements should land as:
- Customer repo + Terraform + CI
- **No** requirement that customer installs platformctl
You use platformctl/wizard on your side to generate PRs; their CI only runs `terraform`. Install platformctl in customer env only if *their* engineers will run discover/wizard themselves.
---
## After apply: incomplete intake, late changes, edit vs teardown
Requirements are never complete on day one. The wizard is **not** a one-shot form you fill once and throw away. It is a **re-entrant way to change intent** over the life of the customer project.
### Mental model
```text
Apply creates managed state
→ later: change intent / decisions (normal path)
→ plan shows create / update / destroy for the delta only
→ apply the delta
```
Terraform already knows how to add, change, and remove resources from a new desired config. The products job is to make those intent edits safe and explainable — not to re-run “full setup from scratch.”
### What you do when something is added or changed later
| Change type | Recourse | Typical commands |
|---|---|---|
| Add domain / env / warehouse profile | **Edit** intent → plan → apply | `wizard` (partial) or hand-edit `intent.yaml` |
| Tighten/loosen access | **Edit** intent + maybe new **decision** → plan → apply | same |
| Reverse a temporary exception | Update/expire **decision**, change override → plan → apply | may **destroy** specific grants/objects |
| Drop a domain / decommission env | **Edit** intent (remove) → plan shows destroys → apply with review | selective teardown via TF |
| Wrong blueprint choice early on | Change intent params or bump blueprint pin → plan (can be large) | prefer migrate path over wipe |
| Abandoned experiment / full reset | Rare **teardown** of a managed stack or env | `platformctl destroy --target …` with guards |
**Default recourse after apply = edit + plan + apply.**
**Teardown = explicit, scoped destroy**, not the normal way to “fix a mistake.”
### How the wizard helps when intake is partial
1. **Partial / resume** — Wizard tracks answered vs unanswered; you can run it again weeks later and only fill new questions (e.g. “SSO ready now?”, “Add marketing domain?”).
2. **Change commands, not only init** — Same catalog supports intents like:
- `wizard --resume`
- `wizard add domain`
- `wizard set identity sso`
- `wizard adopt` (brownfield leftovers)
3. **Diff-aware** — After answers, it patches `intent.yaml` / `decisions/` and runs `plan` so you see *only* the delta against whats already applied.
4. **Decisions for late surprises** — When a late requirement breaks a default, wizard forces a decision record (rationale + optional expiry) instead of a silent HCL tweak.
5. **Drift before change**`drift` first if humans may have clicked in the Snowflake UI; reconcile, then edit intent.
So wizard = **guided editor of intent over time**, not a day-0 questionnaire.
### Edit path (happy path for “things changed”)
```text
Day 0: intent has 2 domains → apply
Day 45: business wants finance + SSO
→ wizard resume / edit intent
→ new decision if needed
→ plan: +finance objects, +SSO stubs, maybe grant changes
→ apply
Day 90: remove temporary finance prod-read
→ expire decision, remove override
→ plan: revoke grants / drop exception role mapping
→ apply
```
No full rebuild. State and decisions carry forward.
### Teardown path (when you actually need destroy)
Use sparingly, always scoped:
- **Resource-level:** remove from intent → TF plans `destroy` for those addresses only
- **Env-level:** `platformctl destroy --env dev` (wrapper around TF destroy targeting that module)
- **Customer-level:** last resort; requires explicit confirmation + maybe separate break-glass role
Guards:
- Policy: destroy of prod requires extra approval / decision
- Refusing destroy when unmanaged dependents exist (warn from discover)
- Prefer “quarantine” (revoke access, keep data) over drop for data-bearing schemas
### What not to do
- Dont re-run greenfield `init` on an existing customer to “pick up new requirements”
- Dont tear down prod to apply a naming tweak — migrate with plan
- Dont use the LLM to emit ad-hoc DROP scripts outside the intent→plan loop
### Short answer
**After something is applied, the recourse is almost always edit (intent/decisions) → plan → apply.**
Wizard helps by making those later edits guided and resumable.
**Teardown is the exception** — selective destroy when you intentionally remove a domain/env/stack, still driven by intent and Terraform, not a separate wipe tool as the main workflow.
---
## Reducing adoption friction (Snowflake people, not “platformctl experts”)
The tool is useful only if a Snowflake developer/devops/admin can use it **without treating it as a second career**. Friction reduction is a first-class design goal, not docs afterthought.
### Design principle
**Speak Snowflake; hide the toolkit.**
Users should think in domains, roles, warehouses, envs — not in blueprint schema versions and planner internals.
### Friction reducers (concrete)
1. **Tiny command surface** — teach 5 verbs, not a platform:
- `wizard` / `add` / `plan` / `apply` / `drift`
Everything else is advanced (`decision`, `adopt`, `destroy`, `blueprint upgrade`).
2. **Snowflake vocabulary in the wizard** — questions use their words (“Who can read curated in prod?”), not ours (“Set access_profile on domain intent key”). Mapping to intent is invisible.
3. **LLM/Cursor skill as the default front door** — for many users the “training” is:
- Open customer repo
- Say: “Add a finance domain with restricted prod access”
- Skill runs validate → plan → shows TF/Snowflake diff in plain language
They never memorize flags. CLI remains for CI and power users.
4. **Show artifacts they already trust** — every plan prints:
- Human summary (“creates 2 roles, 1 warehouse, grants SELECT on FINANCE_CURATED”)
- Familiar `terraform plan`
- Optional SQL preview
No proprietary plan language to learn.
5. **Golden-path recipes, not a manual** — short how-tos aligned to jobs:
- How to add a domain
- How to grant analyst read
- How to add an env
- How to fix drift
Each is 58 steps. Persona = Snowflake admin, not “platformctl operator.”
6. **Sensible defaults everywhere** — first `wizard` run should succeed with mostly Enter/defaults. Customization is opt-in; decisions only when they break a default.
7. **Dont force the tool into daily CI life** — customer pipelines keep running plain Terraform. Only operators who change intent need wizard/CLI/skill. Admins who only approve PRs review the same TF diff they already know.
8. **Escape hatch without shame** — advanced users can edit Terraform modules directly; next `drift`/`plan` still works. The tool must not punish leaving the happy path (document “supported vs at-your-own-risk” edits).
9. **In-repo onboarding, zero slide deck required**:
- `platformctl doctor` — checks auth, backend, blueprint pin, suggests next command
- Customer repo `README` auto-generated: “Your next step is …”
- `platformctl explain` — why this grant exists (reads decisions)
10. **Progressive disclosure** — day 1: wizard + apply. Week 4: decisions. Month 3: blueprint upgrade / adopt. Dont dump the layer cake in training.
### What “trained” should mean (target)
| Role | Must learn | Need not learn |
|---|---|---|
| Snowflake admin | Approve plan summaries; answer wizard in Snowflake terms | Blueprint internals, policy engine |
| DevOps | Wire secrets/state once; TF plan/apply in CI | Wizard catalogs, intent schema |
| Platform operator (you) | Intent, decisions, blueprint pins | — (this is the deep skill) |
If everyone must learn what only operators need, friction has failed.
### Anti-patterns that create overhead
- Requiring a multi-day course before first apply
- Inventing new jargon (`intent compiler`, `L3 model`) in the UI
- Forcing platformctl in every customer engineers toolchain
- Plans that only make sense if you know module graph internals
- Blocking work when someone hand-fixed one grant in Snowflake (prefer detect + guided reconcile)
### Success metric
A Snowflake-competent person who has never seen the tool completes **“add a domain to an existing customer”** in one sitting with only the skill/README — no workshop.
---
## How skills / LLM connect to the tool
Skills are **not a second system**. They are an optional **front door** onto the same customer repo + `platformctl` control plane. The LLM does not apply Snowflake changes by itself; it drives the same files and commands a human would.
### Connection in one picture
```mermaid
flowchart TB
user[User_in_Cursor_or_chat]
skill[Skill_pack_from_blueprint]
ctx[Load_customer_context]
llm[LLM_reasoning]
files[Edit_intent_decisions]
cli[platformctl_validate_plan_drift]
tf[Terraform_apply]
sf[Snowflake]
user --> skill
skill --> ctx
ctx --> llm
skill --> llm
llm -->|"propose_or_write"| files
llm -->|"invoke"| cli
files --> cli
cli --> tf --> sf
```
| Piece | Role |
|---|---|
| **Blueprint `skills/`** | Versioned instructions + best-practice text the agent must load (how we do RBAC, when to require a decision, forbidden patterns) |
| **Customer repo** | Durable context: `intent.yaml`, `decisions/`, TF state ref, last plan — replaces “empty chat memory” |
| **LLM** | Interprets natural language; maps Snowflake-ish requests onto intent edits / wizard answers / CLI calls |
| **`platformctl`** | Source of truth for validate / plan / drift / apply wrappers — deterministic |
| **Terraform** | Still the only applier to Snowflake |
### What a skill actually contains
Shipped inside the blueprint package (pinned with the customers blueprint version), e.g.:
```text
skills/platform-ops/
SKILL.md # when to use, workflow, allowed tools
rbac-best-practices.md # your standards (prose the model follows)
examples.md # few-shot: “add domain” → intent diff
```
`SKILL.md` tells the agent roughly:
1. Open / read `intent.yaml`, `decisions/`, blueprint pin
2. Prefer `platformctl` over inventing SQL/HCL
3. If the request breaks a default → create/update a decision
4. Always run `platformctl plan` before suggesting apply
5. Never apply without explicit user approval
Same catalog the CLI wizard uses can be referenced so chat and CLI stay aligned.
### What the user experiences
**Without skill (CLI-only):**
```bash
platformctl wizard add domain
platformctl plan
platformctl apply
```
**With skill (LLM front door):**
> “Add a finance domain with restricted prod curated read for month-end; revisit after SSO.”
Agent:
1. Loads skill + customer intent/decisions
2. Patches `intent.yaml` (add finance, restricted profile)
3. Writes `decisions/…-finance-prod-read` with rationale + expiry
4. Runs `platformctl validate` + `platformctl plan`
5. Shows summary + terraform plan
6. Waits for you to say apply → `platformctl apply` (or CI PR)
Same end state as CLI; less tool training.
### What the LLM is *not* connected as
- Not a free-form Snowflake admin that emits `GRANT` scripts as the system of record
- Not the store of decisions (files are)
- Not required in customer CI (CI runs terraform; skills stay in operator IDE)
- Not allowed to skip policy gates — `platformctl plan` still fails if override lacks decision
### Why this fixes todays LLM-driven workflow
| Today | With this connection |
|---|---|
| Chat + TF state; why lives in scrollback | Chat + skill + **intent/decisions files** |
| Each session re-derives process | Skill encodes process; blueprint version pins it |
| LLM may invent one-off HCL/SQL | LLM edits intent; planner/TF emit HCL |
| Best practices in peoples heads | Best practices in versioned `skills/` + `policies/` |
### Where skills live vs where they run
| Artifact | Lives in | Runs in |
|---|---|---|
| Skill packs | Blueprint package `skills/` | Operator IDE (Cursor) or chat host |
| `platformctl` | pipx/uv install | Laptop or CI (optional) |
| Apply | Terraform | Customer CI / operator with creds |
Customer engineers who only approve PRs never need the skill. Operators who change platforms use skill **or** CLI — both write the same repo.
### Short definition
**Skills/LLM = natural-language driver for `platformctl` + customer intent/decision files, constrained by blueprint skill packs and policies. Connection point is the customer repo and CLI — not a parallel AI path into Snowflake.**
### Where you type that sentence (Cursor / Claude Code / etc.)
Yes — the natural-language line is typed in an **agent coding tool** such as **Cursor Agent** or **Claude Code** (or similar), with the **customer project repo open as the workspace**. It is not a separate SaaS chat bound to Snowflake.
**Do users need to hunt for skill files?**
Ideally **no**. Skills are installed where the agent auto-discovers them. The user just opens the customer repo and asks in plain language. Looking up `SKILL.md` paths is for authors/operators maintaining the blueprint — not for every Snowflake admin.
#### Recommended layout (Cursor)
| Skill placement | Path | Who gets it |
|---|---|---|
| **Project skill (preferred for delivery)** | `customers/acme/.cursor/skills/platform-ops/SKILL.md` | Anyone who clones the customer repo in Cursor |
| Personal skill (your laptop) | `~/.cursor/skills/platform-ops/` | Only you, across projects |
`platformctl init` / wizard should **copy or link** the blueprints `skills/platform-ops` into the customer repos `.cursor/skills/` so discovery is automatic. Pin the skill content to the same blueprint version as `intent.yaml`.
Claude Code: same idea — ship the skill/instructions into the repo (e.g. project skill / `CLAUDE.md` pointer) so opening the repo is enough; exact folder conventions follow that tools docs.
#### Where to run
| Requirement | Why |
|---|---|
| Workspace root = **customer repo** | Agent sees `intent.yaml`, `decisions/`, Terraform |
| `platformctl` on PATH (operator machine) | Skill invokes validate/plan instead of inventing SQL |
| Snowflake/TF creds as you already use for that customer | Apply/plan still use normal auth — skill doesnt replace that |
| Network to private blueprint registry only if modules arent vendored | Same as non-LLM workflow |
Users do **not** need to open the blueprint monorepo to chat. They open **Acmes customer repo** and talk.
#### Minimal user checklist
1. Clone/open customer repo in Cursor (or Claude Code)
2. Ensure project skill is present (shipped by `init`; if missing, `platformctl skills install`)
3. Ask: “Add finance with restricted prod curated read…”
4. Review plan → approve apply / PR
No manual “find and @-mention the skill file” in the happy path — the description/frontmatter should make the agent pick it up when the request matches (platform change, domain, RBAC, etc.). Power users can still `@` the skill if the host supports explicit attach.
#### What if someone doesnt use Cursor?
They use the **CLI wizard** only — same repo, same outcomes, no skill files involved. Skills are an accelerator for agent IDEs, not a hard dependency of the platform tool.
---
## End-product visualization: new customer scenario
### What you actually open day-to-day
Not a magical Snowflake UI. An **internal toolkit** that looks like:
- A **versioned blueprint repo** (your standard platform modules + policies + skills)
- Per customer: a **customer project repo** containing:
- Terraform (executor)
- `intent.yaml` (what this customer should look like)
- `decisions/` (why deviations exist)
- generated/updated TF from the planner
- A **CLI** (and optionally Cursor skills) you run: `init`, `discover`, `plan`, `apply`, `drift`, `decision`
The LLM is how you talk to that toolkit — it does not replace Terraform.
---
### Scenario A — Mostly greenfield customer (“Acme”)
**Day 0 — Kickoff facts you collect (3060 min)**
You learn: 3 domains (sales, finance, ops), need dev/test/prod, Okta SSO coming later, finance wants tighter access, they already created an empty Snowflake account.
**Step 1 — Create customer project**
```bash
platformctl init acme --blueprint v3
```
Creates repo/folder:
```text
customers/acme/
intent.yaml # mostly empty template
decisions/
terraform/ # wired to blueprint modules
README.md
```
**Step 2 — Fill intent (you or LLM interview)**
You (or chat skill) produce something like:
```yaml
customer: acme
blueprint: v3
environments: [dev, test, prod]
domains:
- name: sales
- name: finance
access_profile: restricted
- name: ops
warehouses:
profile: standard_cost_saver
identity:
sso: planned # not ready yet
```
**Step 3 — Capture a decision where they deviate**
Finance insists analysts can see prod curated early (against your default).
```bash
platformctl decision add \
--applies-to domains.finance \
--rationale "Business needs prod curated read for month-end; revisit after SSO" \
--expires 2026-10-01
```
That writes `decisions/2026-07-15-finance-prod-read.md` (or YAML) linked to the intent override — this is what future-you needed and TF state never had.
**Step 4 — Plan (deterministic + optional LLM explanation)**
```bash
platformctl plan
```
Tool expands blueprint + intent + decisions → Terraform changes:
- env databases/schemas/zones
- role hierarchy + grants
- warehouses
- service users stubs
- finance override applied only where decision allows
You get a normal `terraform plan` style diff + a human summary. Policy checks fail if you skipped a required decision for an override.
**Step 5 — Apply**
```bash
platformctl apply # wraps terraform apply with your CI/approvals
```
Snowflake now matches the plan. TF state is updated as usual.
**Step 6 — Week 3 follow-up (new chat session)**
Someone says: “Add a marketing domain and tighten finance again.”
New LLM session loads **customer repo** (intent + decisions + TF state), not tribal memory.
```bash
platformctl drift # confirms account still matches managed intent
# edit intent: add marketing
platformctl plan
platformctl apply
```
The agent can explain: finance prod-read is still an **active dated decision**, not invent a new story.
---
### Scenario B — Brownfield customer (“Globex”) already has stuff
**Day 0** — They have databases, messy roles, some warehouses, half-manual grants.
**Step 1 — Init + discover**
```bash
platformctl init globex --blueprint v3
platformctl discover --connection globex-prod
```
Output (conceptual):
```text
Found 12 databases, 40 roles, 8 warehouses
Matched to blueprint patterns: 35%
Unmanaged live objects: 62
Conflicts: 4 (naming / dual grant paths)
```
**Step 2 — Adopt, dont rebuild**
Tool proposes an `intent.yaml` draft + `unmanaged.yaml` for things you wont own yet.
You choose:
- Manage: raw/curated zones for 2 domains, new RBAC hierarchy going forward
- Leave unmanaged: legacy `FINANCE_OLD` role until migration date
- Record decision: why legacy role stays
**Step 3 — Import into Terraform**
```bash
platformctl adopt --approve
```
Imports selected objects into TF state / modules so **going forward** changes are planned, not hand-SQL.
**Step 4 — Ongoing**
Same as Acme: intent edits → plan → apply → drift. Brownfield just started with discover/adopt instead of empty init.
---
### Scenario C — What today feels like vs what this product changes
| Moment | Today (LLM + TF) | With this product |
|---|---|---|
| New customer start | Chat invents structure; copy patterns from last engagement | `init` + blueprint + intent template |
| Why finance is special | In someones head / old chat | `decisions/` file loaded every session |
| Follow-up in 2 months | Pass TF state; re-derive intent from HCL | Pass repo; intent+decisions are explicit |
| Client deviation | One-off HCL tweak | Override + required decision + policy gate |
| Brownfield | Manual inventory in chat | `discover` → adopt/unmanaged split |
| Execution | Terraform | Still Terraform |
---
### One-sentence product definition
**An internal CLI + customer-repo convention on top of your Terraform blueprint that stores customer intent and decisions, plans/applies through IaC, discovers brownfield drift, and lets an LLM drive those files instead of reinventing the platform each chat.**