--- name: Snowflake Bootstrap Strategy overview: Internal blueprint repo (modules, schemas, wizard, policies, skills) plus per-customer repos. Wizard drives greenfield/brownfield intent. Customer env gets repo+CI+secrets+Snowflake automation user; blueprint IP stays pinned on operator side by default. todos: - id: blueprint-repo-layout content: Finalize blueprint repo contents (modules, schemas, wizard catalogs, policies, skills, customer-repo template) status: in_progress - id: wizard-catalogs content: Design greenfield and brownfield wizard question catalogs and answer→intent/decision mapping status: pending - id: customer-landing content: Define customer-env bring-up (Snowflake service user, secrets, TF state, git/CI) and default delivery mode A status: pending - id: cli-packaging content: Package platformctl as Python (pipx/uv); publish private; CLI never depends on live access to operator blueprint working tree status: pending - id: lifecycle-edit content: Define post-apply lifecycle — resumable wizard, intent edit→plan→apply as default; scoped destroy as exception status: pending - id: low-friction-ux content: Design low-friction UX — 5 commands, Snowflake vocabulary, LLM skill front door, TF-native plans, recipes, doctor/explain status: pending - id: intent-model content: Define intent/desired-config model above raw TF (blueprint params + customer deviations) status: pending - id: decision-store content: Design decision/rationale store linked to intent paths and overrides status: pending - id: observe-reconcile content: Design observe path — TF state + live Snowflake inventory + drift classification status: pending - id: deep-dive-next content: Next deep-dive — wizard catalog fields vs intent schema (so questions map cleanly) status: pending isProject: false --- # Snowflake Platform Tool: Layered Design (Revised) ## Corrections to prior framing Two important clarifications from you: 1. **Session N+1 is not empty.** You pass Terraform state (and related IaC), so the LLM can reconstruct *what exists*. What does **not** travel with state is **why** a choice was made (rationale, rejected alternatives, temporary compromises, client constraints). 2. **“LLM-driven” ≠ “LLM owns the work.”** Execution and ongoing management are already (or should remain) **IaC / tools**. The LLM is how developers *drive* generation and evolution of that code today — with weak durable intent/decision memory between sessions. Prior plan over-indexed on “empty context” and under-specified the real stack. This revision treats the problem as a **multi-layer system**, not a single bootstrapper slogan. --- ## Precise problem statement | What you already have (or can pass in) | What decays between sessions | |---|---| | TF state / config (what is managed) | Decision rationale (why this shape) | | Live Snowflake (what exists) | Intent vs exception (is this deviation permanent?) | | LLM + developer loop (how changes get authored) | Stable operating procedures as *bound* skills over a schema | | Best practices in people’s heads / docs | Machine-checkable policy + linked ADRs | So the product is not “replace Terraform.” It is: **an intent + decision + policy layer on top of IaC**, with LLM as a skilled driver that must read/write those layers — not only TF state. --- ## Layer stack (this is the real work) ```mermaid flowchart TB subgraph L7 [L7_OperatorUX] chat[LLM_Chat_Skills] cli[CLI_PR_Workflow] end subgraph L6 [L6_PolicyVerify] policy[PolicyAsCode] verify[PostApplyAssertions] end subgraph L5 [L5_PlannerCodegen] planner[IntentDelta_to_TFPlan] codegen[ModuleAware_Codegen] end subgraph L4 [L4_DecisionKnowledge] adr[DecisionRecords] skills[OpsSkills_BestPractices] end subgraph L3 [L3_IntentModel] blueprint[Blueprint_vN] intent[CustomerIntent] overrides[Overrides] end subgraph L2 [L2_Observe] tfstate[TerraformState] inventory[LiveSnowflakeInventory] drift[DriftClassifier] end subgraph L1 [L1_Execute] tf[Terraform_Apply] snow[SnowflakeAccount] end chat --> intent chat --> adr chat --> planner cli --> planner skills --> chat blueprint --> intent overrides --> intent intent --> planner adr --> planner tfstate --> drift inventory --> drift drift --> planner planner --> codegen --> tf --> snow tf --> tfstate snow --> inventory policy --> planner verify --> snow ``` Each layer has its own schema, APIs, and failure modes. They should be designed deliberately — not collapsed into “a bootstrap tool.” --- ### L1 — Execution (already largely exists) **Owns:** applying changes to Snowflake **Artifacts:** Terraform modules/resources, providers, backends, CI apply **Non-goals:** remembering why; encoding client intent in a portable form Assumption going forward: **Terraform remains the executor** (you already pass TF state into sessions). Other engines only matter if you later choose them; do not reopen that as the center of the product. --- ### L2 — Observe / reconcile **Owns:** truthful picture of managed vs live vs drifted Inputs: - Terraform state + config - Live account inventory (SHOW / INFORMATION_SCHEMA / account usage as needed) Outputs: - Resource inventory joined to TF addresses - Drift classes: `in_sync`, `drifted_managed`, `unmanaged_live`, `in_state_missing_live` - Adoption candidates for brownfield This layer answers “where things are.” You already approximate it by feeding TF state to the LLM; inventoriing live Snowflake closes blind spots TF does not see (manual grants, console clicks, out-of-band objects). --- ### L3 — Intent model (above raw TF) **Owns:** *what we mean to run for this customer*, parameterized Artifacts (illustrative): - `blueprint.lock` — pinned platform blueprint version - `intent.yaml` — domains, envs, data products, role assignments, warehouse profiles - `overrides.yaml` — explicit deviations from blueprint defaults Raw TF is a **compiled artifact** of intent. Today the LLM often edits TF directly; that works, but intent stays implicit inside HCL. Making intent explicit is what makes “tweak per client” repeatable without re-deriving the whole graph from chat. Deviations live here as data, not as undocumented HCL quirks. --- ### L4 — Decision / knowledge (the real N+1 gap) **Owns:** *why* — durable across sessions even when TF state is provided Artifacts: - Decision records linked to resources or intent keys, e.g.: - `decision_id`, `applies_to` (TF address or intent path), `status` (active/superseded), `rationale`, `alternatives_rejected`, `expires_on`, `client_constraint` - Curated **ops skills / best-practice packs** (versioned markdown/YAML the LLM must load) - Mapping: blueprint rule → allowed override types → required decision fields Example of what TF state cannot tell the next session: - “We kept a flat role for finance because their IdP groups land weekly; revisit after SCIM.” - “Prod shares DEV database naming temporarily for migration; remove after cutover date.” - “Blueprint says account-per-env; this client insisted on database-per-env for cost.” Without L4, every N+1 session can rebuild *structure* from state and still **re-litigate or contradict** prior judgment. --- ### L5 — Planner / codegen **Owns:** turning intent deltas (+ decisions) into concrete IaC changes Flow: 1. Load blueprint + intent + overrides + decisions 2. Load observe/drift report 3. Compute target module graph / resource set 4. Emit TF changes (or a plan preview) — **schema-validated** 5. Attach/update decision records for any override touched LLM role here: assist codegen and explain diffs **inside this pipeline**, not free-form “write some Snowflake SQL.” Guardrails: module allowlist, forbidden patterns, required decision when overriding policy. --- ### L6 — Policy / verify **Owns:** machine-checkable best practices (not only prose skills) Examples: - No `ACCOUNTADMIN` for service users - Raw zone writable only by loader roles - Env isolation invariants - Naming / tagging requirements - Override must reference an active decision record Runs on: pre-plan validate, PR checks, post-apply assertions. Skills (L4) teach the LLM; policy (L6) **fails the build** when the LLM or a human skips the rules. --- ### L7 — Operator UX **Owns:** how developers drive the system - CLI: `discover`, `intent validate`, `plan`, `apply`, `drift`, `decision add` - LLM skills: load L3+L4+L2 context automatically; propose intent/override/decision patches; call planner - PR workflow: intent/decision/TF in one reviewable unit This is where “LLM-driven” correctly lives — as the **driver**, with IaC as the **engine**. --- ## What “repeatability with deviation” means in this stack ```text Blueprint_vN + Customer intent (params) + Overrides (structured deviations) + Decisions (why those overrides exist) → Planner → Terraform → Snowflake → Observe → Drift → next change ``` Same steps, different inputs. Not the same HCL pasted every time; not a blank-slate LLM redesign every time either. --- ## Implementation breadth (why this feels “many layers”) Rough build slices (can be separate workstreams): 1. **Customer project layout** — conventions for TF root, intent, decisions, skills pin 2. **Blueprint modules** — your recommended infra as versioned TF modules (not one-off roots) 3. **Intent schema + compiler** — intent → module inputs (deterministic first) 4. **Decision store format + linking** — resource/intent annotations 5. **Discover/drift** — state + live inventory join 6. **Policy pack** — OPA/Conftest/custom validators on intent+TF 7. **LLM skill pack** — Cursor/agent skills that require reading L2–L4 before editing 8. **Brownfield adopt** — import live → intent proposal → unmanaged list 9. **Verification suite** — Snowflake-side integration checks per customer 10. **Multi-customer packaging** — how blueprints version and how clients pin upgrades This is product/platform engineering, not a weekend script. The earlier “bootstrapper” framing was the tip of that iceberg. --- ## Recommended discussion order (next deep-dives) Do not design all layers at equal depth at once. Suggested sequence: 1. **L3 Intent model** — what must be expressible without dropping to raw TF 2. **L4 Decision store** — schema that captures the why you lose today 3. **L5 Planner contract** — how intent patches become module-safe TF diffs 4. **L2 Observe** — drift model given you already use TF state 5. **L6 Policy** — which best practices become hard gates first 6. **L7 Skills** — only after schemas exist so skills have something binding to edit --- ## Verdict (revised) - You are already closer than “LLM owns infra”: **IaC owns execution; LLM drives authoring.** - Passing TF state into session N+1 solves **topology memory**, not **decision memory**. - The useful internal product is a **layered intent/decision/policy system over Terraform**, with discover/drift and LLM skills as interfaces — not a single high-level “bootstrap tool” narrative. - Next value is to **deep-dive one layer at a time** and nail schemas/contracts before writing a lot of code. --- ## Blueprint repo — what lives there The blueprint repo is **your IP / product core**. Customer repos consume a pinned version of it; they do not fork a copy of everything. ```text snowflake-platform-blueprint/ modules/ # Terraform modules (executor building blocks) environment/ # env isolation pattern database_zones/ # landing/raw/curated/marts (or your real zones) rbac/ # role hierarchy + grant patterns warehouses/ service_principals/ monitoring/ # resource monitors, basic alerts stubs schemas/ intent.schema.json # what intent.yaml must look like decision.schema.json discover-report.schema.json defaults/ intent.greenfield.yaml # starter intent intent.brownfield.yaml warehouse_profiles.yaml access_profiles.yaml # e.g. standard vs restricted wizard/ catalog.greenfield.yaml # ordered questions, branching, defaults catalog.brownfield.yaml # starts with discover-driven questions mapping.yaml # answer_id → intent path / override / decision prompt policies/ # hard gates (Conftest/OPA/custom) no_accountadmin_services.rego env_isolation.rego override_requires_decision.rego skills/ # LLM operator packs (Cursor/agent skills) platform-ops/SKILL.md rbac-best-practices.md brownfield-adopt.md templates/ customer-repo/ # skeleton copied by `init` intent.yaml decisions/.gitkeep terraform/ .github/workflows/ # optional CI template for customer tests/ reference-customer/ # golden intent → expected plan fixtures CHANGELOG.md # blueprint v1, v2, … VERSION ``` **Rule of thumb** | In blueprint repo | In customer repo | |---|---| | Modules, schemas, wizard catalogs, policies, skills, defaults | Filled `intent.yaml`, `decisions/`, TF root that *calls* modules, state backend config, connection secrets refs | | Versioned releases (`v3.2.0`) | `blueprint: v3.2.0` pin | | Your best-practice brain | This client’s answers and exceptions | --- ## Wizard — yes, first-class The wizard is the primary greenfield/brownfield onboarding UX. It is **not** free-form chat inventing infra; it is a **question catalog** over the blueprint that writes `intent.yaml` + prompts for `decisions/` when answers deviate from defaults. ### Shape ```mermaid flowchart LR mode{Greenfield_or_Brownfield} discover[DiscoverLiveAccount] ask[WizardQuestionCatalog] intent[Write_intent.yaml] decide[PromptDecisionIfOverride] plan[platformctl_plan] apply[ApplyInAllowedContext] mode -->|greenfield| ask mode -->|brownfield| discover --> ask ask --> intent --> decide --> plan --> apply ``` ### Greenfield wizard (examples) Ordered questions from `wizard/catalog.greenfield.yaml`: 1. Customer name / account identifiers 2. Env model: account-per-env vs db-per-env (default from blueprint) 3. Domains / data products (repeatable) 4. Access profile per domain (standard / restricted) 5. Warehouse profile 6. Identity: SSO now / later / service users only 7. Anything that contradicts a default → **forced decision capture** (rationale + optional expiry) Answers compile via `wizard/mapping.yaml` → `intent.yaml`. CLI shape: ```bash platformctl wizard --customer acme --mode greenfield # or non-interactive from a saved answers file: platformctl wizard --answers acme-answers.yaml ``` LLM can *run* the same catalog conversationally (“ask next unanswered question”), but the catalog is the contract — same outputs whether human CLI or chat. ### Brownfield wizard 1. Connect + `discover` 2. Show match % / unmanaged / conflicts 3. Questions become **adoption choices**: manage / ignore / migrate-later per object group 4. Propose intent draft from what matched 5. Require decisions for keep-legacy and naming conflicts 6. `adopt` → import selected into TF → plan residual gaps Same wizard engine; different catalog + discover as step 0. --- ## Bringing this into a new customer environment Split **operator toolkit** (yours) from **runtime footprint** (what must exist in the customer’s world). ### Default delivery model (recommended) **You keep the blueprint + wizard.** The customer receives a **customer project** that can plan/apply against *their* Snowflake, usually from *their* Git + CI — without needing your full blueprint source if modules are consumed as a versioned package (git tag, private module registry, or vendored lock). ```mermaid flowchart TB subgraph your_side [YourSide_Operator] blueprint[BlueprintRepo] wizard[Wizard_CLI_or_Chat] end subgraph customer_side [CustomerEnvironment] crepo[CustomerGitRepo] ci[CustomerCI] tfstate[TFStateBackend] sf[SnowflakeAccount] end wizard -->|generates_or_updates| crepo blueprint -->|pinned_modules| crepo crepo --> ci --> tfstate ci -->|terraform_apply| sf ``` **What must be set up once in the customer environment** 1. **Snowflake bootstrap identity** — service user + key pair (or OAuth), roles enough for Terraform (often a dedicated automation role; initial ACCOUNTADMIN bootstrap is one-time manual) 2. **Secrets** — private key / cloud creds in customer secret store (GitHub Actions secrets, Vault, etc.) — never in git 3. **TF state backend** — customer-controlled (S3/Azure/GCS + lock), or your org’s backend if *you* operate the account under contract 4. **Customer git repo** — the generated project from `init` / wizard 5. **CI workflow** — plan on PR, apply on merge/approval (template from blueprint `templates/customer-repo`) **What does *not* need to live in the customer env** - Wizard question catalogs (can run on your laptop during engagement) - Full policy authoring / skill authoring - Blueprint module *source* (if you publish versioned modules they consume read-only) ### How day-1 “bring-up” works in practice 1. Customer creates (or you create) empty Snowflake account / grants you access 2. One-time manual: create Terraform service user + store key in their secrets 3. You run wizard (locally or internal portal) → produces customer repo contents 4. Push customer repo to **their** GitHub/GitLab (or your managed repo if contract says so) 5. Wire secrets + state backend 6. CI runs `plan`/`apply` into their account 7. Ongoing changes: PR to customer repo (intent/decision/TF); wizard/LLM used by your team as operators ### Alternate operating modes (when needed) | Mode | When | Footprint in customer | |---|---|---| | **A. Customer-hosted CI (default)** | Normal consulting delivery, customer owns account | Repo + secrets + state + Snowflake service user | | **B. You operate apply** | You manage platform under retainer | Same Snowflake identity; state/CI may stay in your org | | **C. Fully offline customer** | No external module fetch | Vendor/pin blueprint modules into customer repo at release time | Avoid shipping the entire blueprint monorepo into every customer unless they require air-gapped ownership — that blurs IP and versioning. ### “Tool in customer environment” checklist (handoff) - [ ] Snowflake automation user created - [ ] Secrets installed in CI - [ ] State backend ready - [ ] Customer repo pushed with blueprint pin - [ ] First successful `plan` in CI - [ ] First successful `apply` - [ ] Drift job scheduled or documented - [ ] Runbook: how to add a domain / raise a decision / upgrade blueprint pin --- ## CLI access model and packaging ### Will the CLI reach into “my” repo from the customer environment? **No — not if designed correctly.** That would be brittle (VPN, private git auth, IP leakage) and is the wrong dependency direction. Correct dependency direction: ```text Published artifacts (versioned) ├── platformctl (CLI package on private PyPI / uv tool) └── blueprint modules + schemas + wizard catalogs (versioned release, e.g. v3.2.0) Customer environment / CI only needs: ├── customer git repo (intent, decisions, TF root) ├── platformctl (optional; CI may only need terraform) └── ability to fetch pinned blueprint vX (registry / git tag / vendored copy) ``` The CLI works against the **local customer project directory** (and Snowflake + TF state). It does **not** SSH into your laptop or clone your live blueprint working tree. | Who runs what | Accesses | |---|---| | You on laptop (wizard during engagement) | Local checkout or cached blueprint package + writes customer repo | | Customer CI `terraform plan/apply` | Customer repo + pinned TF modules only (often **no** platformctl) | | Customer engineer with platformctl (optional) | Same: local customer repo + installed CLI + pinned blueprint package from registry | If the customer is air-gapped: **vendor** blueprint `v3.2.0` into the customer repo at release time (mode C). Still no live link to your private monorepo. ### Packaging recommendation **Default: Python package, installed via `pipx` or `uv tool install`.** Why Python over npm/Go for this tool: - Snowflake ecosystem is Python-native (`snowflake-connector-python`, many internal data-platform scripts) - Excellent YAML/JSON schema, Jinja, Click/Typer CLIs - Easy private publish (Azure Artifacts, AWS CodeArtifact, GitHub Packages, private PyPI, or `pip install git+https://…`) - `pipx`/`uvx` gives an isolated CLI without polluting customer project venvs | Option | Fit | Use when | |---|---|---| | **Python + pipx/uv (recommended)** | Best overall | Your team already does Snowflake/Python | | **Go single binary** | Best “download one file” DX | Strict customer lockdown, no Python allowed on jump hosts | | **npm** | Weak fit | Only if your org is JS-only; TF/Snowflake glue will feel awkward | | **Snowflake CLI plugin** | Niche | If you standardize on `snow` CLI everywhere; still thin wrapper over same Python core | | **Docker image** | Good CI companion | `ghcr.io/yourorg/platformctl:3.2.0` for customer pipelines without local install | Suggested package split: 1. **`platformctl`** — CLI (wizard, validate, plan wrapper, decision helpers, discover) 2. **`platform-blueprint`** (or git-tagged module repo) — versioned content the CLI resolves by `blueprint: v3.2.0` in intent Publish both to **your** private registry. Customer CI authenticates to that registry *or* you vendor the blueprint pin into their repo so they need zero access to your org. ### Minimal customer footprint (preferred) Many engagements should land as: - Customer repo + Terraform + CI - **No** requirement that customer installs platformctl You use platformctl/wizard on your side to generate PRs; their CI only runs `terraform`. Install platformctl in customer env only if *their* engineers will run discover/wizard themselves. --- ## After apply: incomplete intake, late changes, edit vs teardown Requirements are never complete on day one. The wizard is **not** a one-shot form you fill once and throw away. It is a **re-entrant way to change intent** over the life of the customer project. ### Mental model ```text Apply creates managed state → later: change intent / decisions (normal path) → plan shows create / update / destroy for the delta only → apply the delta ``` Terraform already knows how to add, change, and remove resources from a new desired config. The product’s job is to make those intent edits safe and explainable — not to re-run “full setup from scratch.” ### What you do when something is added or changed later | Change type | Recourse | Typical commands | |---|---|---| | Add domain / env / warehouse profile | **Edit** intent → plan → apply | `wizard` (partial) or hand-edit `intent.yaml` | | Tighten/loosen access | **Edit** intent + maybe new **decision** → plan → apply | same | | Reverse a temporary exception | Update/expire **decision**, change override → plan → apply | may **destroy** specific grants/objects | | Drop a domain / decommission env | **Edit** intent (remove) → plan shows destroys → apply with review | selective teardown via TF | | Wrong blueprint choice early on | Change intent params or bump blueprint pin → plan (can be large) | prefer migrate path over wipe | | Abandoned experiment / full reset | Rare **teardown** of a managed stack or env | `platformctl destroy --target …` with guards | **Default recourse after apply = edit + plan + apply.** **Teardown = explicit, scoped destroy**, not the normal way to “fix a mistake.” ### How the wizard helps when intake is partial 1. **Partial / resume** — Wizard tracks answered vs unanswered; you can run it again weeks later and only fill new questions (e.g. “SSO ready now?”, “Add marketing domain?”). 2. **Change commands, not only init** — Same catalog supports intents like: - `wizard --resume` - `wizard add domain` - `wizard set identity sso` - `wizard adopt` (brownfield leftovers) 3. **Diff-aware** — After answers, it patches `intent.yaml` / `decisions/` and runs `plan` so you see *only* the delta against what’s already applied. 4. **Decisions for late surprises** — When a late requirement breaks a default, wizard forces a decision record (rationale + optional expiry) instead of a silent HCL tweak. 5. **Drift before change** — `drift` first if humans may have clicked in the Snowflake UI; reconcile, then edit intent. So wizard = **guided editor of intent over time**, not a day-0 questionnaire. ### Edit path (happy path for “things changed”) ```text Day 0: intent has 2 domains → apply Day 45: business wants finance + SSO → wizard resume / edit intent → new decision if needed → plan: +finance objects, +SSO stubs, maybe grant changes → apply Day 90: remove temporary finance prod-read → expire decision, remove override → plan: revoke grants / drop exception role mapping → apply ``` No full rebuild. State and decisions carry forward. ### Teardown path (when you actually need destroy) Use sparingly, always scoped: - **Resource-level:** remove from intent → TF plans `destroy` for those addresses only - **Env-level:** `platformctl destroy --env dev` (wrapper around TF destroy targeting that module) - **Customer-level:** last resort; requires explicit confirmation + maybe separate break-glass role Guards: - Policy: destroy of prod requires extra approval / decision - Refusing destroy when unmanaged dependents exist (warn from discover) - Prefer “quarantine” (revoke access, keep data) over drop for data-bearing schemas ### What not to do - Don’t re-run greenfield `init` on an existing customer to “pick up new requirements” - Don’t tear down prod to apply a naming tweak — migrate with plan - Don’t use the LLM to emit ad-hoc DROP scripts outside the intent→plan loop ### Short answer **After something is applied, the recourse is almost always edit (intent/decisions) → plan → apply.** Wizard helps by making those later edits guided and resumable. **Teardown is the exception** — selective destroy when you intentionally remove a domain/env/stack, still driven by intent and Terraform, not a separate wipe tool as the main workflow. --- ## Reducing adoption friction (Snowflake people, not “platformctl experts”) The tool is useful only if a Snowflake developer/devops/admin can use it **without treating it as a second career**. Friction reduction is a first-class design goal, not docs afterthought. ### Design principle **Speak Snowflake; hide the toolkit.** Users should think in domains, roles, warehouses, envs — not in blueprint schema versions and planner internals. ### Friction reducers (concrete) 1. **Tiny command surface** — teach 5 verbs, not a platform: - `wizard` / `add` / `plan` / `apply` / `drift` Everything else is advanced (`decision`, `adopt`, `destroy`, `blueprint upgrade`). 2. **Snowflake vocabulary in the wizard** — questions use their words (“Who can read curated in prod?”), not ours (“Set access_profile on domain intent key”). Mapping to intent is invisible. 3. **LLM/Cursor skill as the default front door** — for many users the “training” is: - Open customer repo - Say: “Add a finance domain with restricted prod access” - Skill runs validate → plan → shows TF/Snowflake diff in plain language They never memorize flags. CLI remains for CI and power users. 4. **Show artifacts they already trust** — every plan prints: - Human summary (“creates 2 roles, 1 warehouse, grants SELECT on FINANCE_CURATED”) - Familiar `terraform plan` - Optional SQL preview No proprietary plan language to learn. 5. **Golden-path recipes, not a manual** — short how-tos aligned to jobs: - How to add a domain - How to grant analyst read - How to add an env - How to fix drift Each is 5–8 steps. Persona = Snowflake admin, not “platformctl operator.” 6. **Sensible defaults everywhere** — first `wizard` run should succeed with mostly Enter/defaults. Customization is opt-in; decisions only when they break a default. 7. **Don’t force the tool into daily CI life** — customer pipelines keep running plain Terraform. Only operators who change intent need wizard/CLI/skill. Admins who only approve PRs review the same TF diff they already know. 8. **Escape hatch without shame** — advanced users can edit Terraform modules directly; next `drift`/`plan` still works. The tool must not punish leaving the happy path (document “supported vs at-your-own-risk” edits). 9. **In-repo onboarding, zero slide deck required**: - `platformctl doctor` — checks auth, backend, blueprint pin, suggests next command - Customer repo `README` auto-generated: “Your next step is …” - `platformctl explain` — why this grant exists (reads decisions) 10. **Progressive disclosure** — day 1: wizard + apply. Week 4: decisions. Month 3: blueprint upgrade / adopt. Don’t dump the layer cake in training. ### What “trained” should mean (target) | Role | Must learn | Need not learn | |---|---|---| | Snowflake admin | Approve plan summaries; answer wizard in Snowflake terms | Blueprint internals, policy engine | | DevOps | Wire secrets/state once; TF plan/apply in CI | Wizard catalogs, intent schema | | Platform operator (you) | Intent, decisions, blueprint pins | — (this is the deep skill) | If everyone must learn what only operators need, friction has failed. ### Anti-patterns that create overhead - Requiring a multi-day course before first apply - Inventing new jargon (`intent compiler`, `L3 model`) in the UI - Forcing platformctl in every customer engineer’s toolchain - Plans that only make sense if you know module graph internals - Blocking work when someone hand-fixed one grant in Snowflake (prefer detect + guided reconcile) ### Success metric A Snowflake-competent person who has never seen the tool completes **“add a domain to an existing customer”** in one sitting with only the skill/README — no workshop. --- ## How skills / LLM connect to the tool Skills are **not a second system**. They are an optional **front door** onto the same customer repo + `platformctl` control plane. The LLM does not apply Snowflake changes by itself; it drives the same files and commands a human would. ### Connection in one picture ```mermaid flowchart TB user[User_in_Cursor_or_chat] skill[Skill_pack_from_blueprint] ctx[Load_customer_context] llm[LLM_reasoning] files[Edit_intent_decisions] cli[platformctl_validate_plan_drift] tf[Terraform_apply] sf[Snowflake] user --> skill skill --> ctx ctx --> llm skill --> llm llm -->|"propose_or_write"| files llm -->|"invoke"| cli files --> cli cli --> tf --> sf ``` | Piece | Role | |---|---| | **Blueprint `skills/`** | Versioned instructions + best-practice text the agent must load (how we do RBAC, when to require a decision, forbidden patterns) | | **Customer repo** | Durable context: `intent.yaml`, `decisions/`, TF state ref, last plan — replaces “empty chat memory” | | **LLM** | Interprets natural language; maps Snowflake-ish requests onto intent edits / wizard answers / CLI calls | | **`platformctl`** | Source of truth for validate / plan / drift / apply wrappers — deterministic | | **Terraform** | Still the only applier to Snowflake | ### What a skill actually contains Shipped inside the blueprint package (pinned with the customer’s blueprint version), e.g.: ```text skills/platform-ops/ SKILL.md # when to use, workflow, allowed tools rbac-best-practices.md # your standards (prose the model follows) examples.md # few-shot: “add domain” → intent diff ``` `SKILL.md` tells the agent roughly: 1. Open / read `intent.yaml`, `decisions/`, blueprint pin 2. Prefer `platformctl` over inventing SQL/HCL 3. If the request breaks a default → create/update a decision 4. Always run `platformctl plan` before suggesting apply 5. Never apply without explicit user approval Same catalog the CLI wizard uses can be referenced so chat and CLI stay aligned. ### What the user experiences **Without skill (CLI-only):** ```bash platformctl wizard add domain platformctl plan platformctl apply ``` **With skill (LLM front door):** > “Add a finance domain with restricted prod curated read for month-end; revisit after SSO.” Agent: 1. Loads skill + customer intent/decisions 2. Patches `intent.yaml` (add finance, restricted profile) 3. Writes `decisions/…-finance-prod-read` with rationale + expiry 4. Runs `platformctl validate` + `platformctl plan` 5. Shows summary + terraform plan 6. Waits for you to say apply → `platformctl apply` (or CI PR) Same end state as CLI; less tool training. ### What the LLM is *not* connected as - Not a free-form Snowflake admin that emits `GRANT` scripts as the system of record - Not the store of decisions (files are) - Not required in customer CI (CI runs terraform; skills stay in operator IDE) - Not allowed to skip policy gates — `platformctl plan` still fails if override lacks decision ### Why this fixes today’s LLM-driven workflow | Today | With this connection | |---|---| | Chat + TF state; why lives in scrollback | Chat + skill + **intent/decisions files** | | Each session re-derives process | Skill encodes process; blueprint version pins it | | LLM may invent one-off HCL/SQL | LLM edits intent; planner/TF emit HCL | | Best practices in people’s heads | Best practices in versioned `skills/` + `policies/` | ### Where skills live vs where they run | Artifact | Lives in | Runs in | |---|---|---| | Skill packs | Blueprint package `skills/` | Operator IDE (Cursor) or chat host | | `platformctl` | pipx/uv install | Laptop or CI (optional) | | Apply | Terraform | Customer CI / operator with creds | Customer engineers who only approve PRs never need the skill. Operators who change platforms use skill **or** CLI — both write the same repo. ### Short definition **Skills/LLM = natural-language driver for `platformctl` + customer intent/decision files, constrained by blueprint skill packs and policies. Connection point is the customer repo and CLI — not a parallel AI path into Snowflake.** ### Where you type that sentence (Cursor / Claude Code / etc.) Yes — the natural-language line is typed in an **agent coding tool** such as **Cursor Agent** or **Claude Code** (or similar), with the **customer project repo open as the workspace**. It is not a separate SaaS chat bound to Snowflake. **Do users need to hunt for skill files?** Ideally **no**. Skills are installed where the agent auto-discovers them. The user just opens the customer repo and asks in plain language. Looking up `SKILL.md` paths is for authors/operators maintaining the blueprint — not for every Snowflake admin. #### Recommended layout (Cursor) | Skill placement | Path | Who gets it | |---|---|---| | **Project skill (preferred for delivery)** | `customers/acme/.cursor/skills/platform-ops/SKILL.md` | Anyone who clones the customer repo in Cursor | | Personal skill (your laptop) | `~/.cursor/skills/platform-ops/` | Only you, across projects | `platformctl init` / wizard should **copy or link** the blueprint’s `skills/platform-ops` into the customer repo’s `.cursor/skills/` so discovery is automatic. Pin the skill content to the same blueprint version as `intent.yaml`. Claude Code: same idea — ship the skill/instructions into the repo (e.g. project skill / `CLAUDE.md` pointer) so opening the repo is enough; exact folder conventions follow that tool’s docs. #### Where to run | Requirement | Why | |---|---| | Workspace root = **customer repo** | Agent sees `intent.yaml`, `decisions/`, Terraform | | `platformctl` on PATH (operator machine) | Skill invokes validate/plan instead of inventing SQL | | Snowflake/TF creds as you already use for that customer | Apply/plan still use normal auth — skill doesn’t replace that | | Network to private blueprint registry only if modules aren’t vendored | Same as non-LLM workflow | Users do **not** need to open the blueprint monorepo to chat. They open **Acme’s customer repo** and talk. #### Minimal user checklist 1. Clone/open customer repo in Cursor (or Claude Code) 2. Ensure project skill is present (shipped by `init`; if missing, `platformctl skills install`) 3. Ask: “Add finance with restricted prod curated read…” 4. Review plan → approve apply / PR No manual “find and @-mention the skill file” in the happy path — the description/frontmatter should make the agent pick it up when the request matches (platform change, domain, RBAC, etc.). Power users can still `@` the skill if the host supports explicit attach. #### What if someone doesn’t use Cursor? They use the **CLI wizard** only — same repo, same outcomes, no skill files involved. Skills are an accelerator for agent IDEs, not a hard dependency of the platform tool. --- ## End-product visualization: new customer scenario ### What you actually open day-to-day Not a magical Snowflake UI. An **internal toolkit** that looks like: - A **versioned blueprint repo** (your standard platform modules + policies + skills) - Per customer: a **customer project repo** containing: - Terraform (executor) - `intent.yaml` (what this customer should look like) - `decisions/` (why deviations exist) - generated/updated TF from the planner - A **CLI** (and optionally Cursor skills) you run: `init`, `discover`, `plan`, `apply`, `drift`, `decision` The LLM is how you talk to that toolkit — it does not replace Terraform. --- ### Scenario A — Mostly greenfield customer (“Acme”) **Day 0 — Kickoff facts you collect (30–60 min)** You learn: 3 domains (sales, finance, ops), need dev/test/prod, Okta SSO coming later, finance wants tighter access, they already created an empty Snowflake account. **Step 1 — Create customer project** ```bash platformctl init acme --blueprint v3 ``` Creates repo/folder: ```text customers/acme/ intent.yaml # mostly empty template decisions/ terraform/ # wired to blueprint modules README.md ``` **Step 2 — Fill intent (you or LLM interview)** You (or chat skill) produce something like: ```yaml customer: acme blueprint: v3 environments: [dev, test, prod] domains: - name: sales - name: finance access_profile: restricted - name: ops warehouses: profile: standard_cost_saver identity: sso: planned # not ready yet ``` **Step 3 — Capture a decision where they deviate** Finance insists analysts can see prod curated early (against your default). ```bash platformctl decision add \ --applies-to domains.finance \ --rationale "Business needs prod curated read for month-end; revisit after SSO" \ --expires 2026-10-01 ``` That writes `decisions/2026-07-15-finance-prod-read.md` (or YAML) linked to the intent override — this is what future-you needed and TF state never had. **Step 4 — Plan (deterministic + optional LLM explanation)** ```bash platformctl plan ``` Tool expands blueprint + intent + decisions → Terraform changes: - env databases/schemas/zones - role hierarchy + grants - warehouses - service users stubs - finance override applied only where decision allows You get a normal `terraform plan` style diff + a human summary. Policy checks fail if you skipped a required decision for an override. **Step 5 — Apply** ```bash platformctl apply # wraps terraform apply with your CI/approvals ``` Snowflake now matches the plan. TF state is updated as usual. **Step 6 — Week 3 follow-up (new chat session)** Someone says: “Add a marketing domain and tighten finance again.” New LLM session loads **customer repo** (intent + decisions + TF state), not tribal memory. ```bash platformctl drift # confirms account still matches managed intent # edit intent: add marketing platformctl plan platformctl apply ``` The agent can explain: finance prod-read is still an **active dated decision**, not invent a new story. --- ### Scenario B — Brownfield customer (“Globex”) already has stuff **Day 0** — They have databases, messy roles, some warehouses, half-manual grants. **Step 1 — Init + discover** ```bash platformctl init globex --blueprint v3 platformctl discover --connection globex-prod ``` Output (conceptual): ```text Found 12 databases, 40 roles, 8 warehouses Matched to blueprint patterns: 35% Unmanaged live objects: 62 Conflicts: 4 (naming / dual grant paths) ``` **Step 2 — Adopt, don’t rebuild** Tool proposes an `intent.yaml` draft + `unmanaged.yaml` for things you won’t own yet. You choose: - Manage: raw/curated zones for 2 domains, new RBAC hierarchy going forward - Leave unmanaged: legacy `FINANCE_OLD` role until migration date - Record decision: why legacy role stays **Step 3 — Import into Terraform** ```bash platformctl adopt --approve ``` Imports selected objects into TF state / modules so **going forward** changes are planned, not hand-SQL. **Step 4 — Ongoing** Same as Acme: intent edits → plan → apply → drift. Brownfield just started with discover/adopt instead of empty init. --- ### Scenario C — What today feels like vs what this product changes | Moment | Today (LLM + TF) | With this product | |---|---|---| | New customer start | Chat invents structure; copy patterns from last engagement | `init` + blueprint + intent template | | Why finance is special | In someone’s head / old chat | `decisions/` file loaded every session | | Follow-up in 2 months | Pass TF state; re-derive intent from HCL | Pass repo; intent+decisions are explicit | | Client deviation | One-off HCL tweak | Override + required decision + policy gate | | Brownfield | Manual inventory in chat | `discover` → adopt/unmanaged split | | Execution | Terraform | Still Terraform | --- ### One-sentence product definition **An internal CLI + customer-repo convention on top of your Terraform blueprint that stores customer intent and decisions, plans/applies through IaC, discovers brownfield drift, and lets an LLM drive those files instead of reinventing the platform each chat.**