# dbt — API: Four Verbs and a Fan-Out: The dbt Integration over HTTP

The dbt connector declares Import, Audit, and Generate, and four routes carry the loop that matters: push artifacts in, check them against governed meaning, publish contracts back out, and ask what the last import knew. This is the route-by-route reference — role gating, request bodies field by field, exact response shapes. Throughout, `https://coremodels.example.com` stands in for your deployment, `$TOKEN` for your bearer token, and `$PROJECT_ID` for the 32-character hex id of the governing project. The vendor key is `dbt`. One property runs through all four: no credentials — artifacts in, artifacts out. CoreModels never connects to your warehouse, never runs dbt, never reads a dbt platform API, and never writes into your repository; Generate hands back files, and your own PR flow lands them.

# Four Verbs and a Fan-Out: The dbt Integration over HTTP

The dbt connector declares Import, Audit, and Generate, and four routes carry the loop that matters: push artifacts in, check them against governed meaning, publish contracts back out, and ask what the last import knew. This is the route-by-route reference — role gating, request bodies field by field, exact response shapes. Throughout, `https://coremodels.example.com` stands in for your deployment, `$TOKEN` for your bearer token, and `$PROJECT_ID` for the 32-character hex id of the governing project. The vendor key is `dbt`. One property runs through all four: no credentials — artifacts in, artifacts out. CoreModels never connects to your warehouse, never runs dbt, never reads a dbt platform API, and never writes into your repository; Generate hands back files, and your own PR flow lands them.

| Route | Method | Role | Writes? |
|---|---|---|---|
| `graph/integrations/dbt/import/{projectId}` | POST | Admin | additive graph writes |
| `graph/integrations/dbt/audit/{projectId}` | POST | Viewer | no (history opt-in) |
| `graph/integrations/dbt/generate/{projectId}` | POST | Viewer | no |
| `graph/integrations/dbt/status/{projectId}` | GET | Viewer | no |

## Import — Admin

```http
POST https://coremodels.example.com/graph/integrations/dbt/import/$PROJECT_ID
Authorization: Bearer $TOKEN
Content-Type: application/json

{ "artifacts": { "manifest": "<target/manifest.json>", "catalog": "<target/catalog.json>" },
  "spaces": [] }
```

`artifacts` maps artifact name to raw file content: `manifest` is required (schema v10 through v12), `catalog` and `semantic_manifest` are optional. `spaces` is an optional array of space ids — empty or omitted means the project's main space. Models, seeds, snapshots, and sources become Types; columns become Elements; `accepted_values` tests become Taxonomies with controlled lists; `relationships` tests become references; `parent_map` becomes Depends On lineage; semantic models and exposures become Components. The response is a counts object — `datasetsAdded`, `datasetsSkippedExisting`, `fieldsAdded`, `lineageEdgesAdded`, `lineageEdgesSkipped`, `nodesEnriched`, `resourcesWritten`, `resourceEdgesAdded`, `snapshotStored` — plus the `lossiness` and `errors` channels. Import is additive: a re-import adds what is new and never mutates governed nodes, so the descriptions, vocabularies, and ontology bindings a human added on top survive every refresh. What changed surfaces through the audit instead.

## Audit — Viewer

```http
POST https://coremodels.example.com/graph/integrations/dbt/audit/$PROJECT_ID

{ "artifacts": { "manifest": "<fresh manifest.json>" }, "recordHistory": true }
```

The same body as import, plus `recordHistory` (default `false` — the audit verb stays strictly read-only unless asked). The response carries `errorCount`, `warningCount`, `infoCount`, a `codes` map of per-code totals, `driftedObjects`, the artifact `fingerprint`, structured `findings`, a `markdown` report ready to paste into a pull request, `historyRecorded`, and `lossiness`. The dbt-specific finding codes are `contract-not-enforced`, `contract-column-missing-type`, `key-column-untested`, and `source-no-freshness`; the shared drift and coverage layer adds `field-type-drift`, `enum-narrowed`, `enum-widened`, `enum-constraint-removed`, and `dataset-removed`. The CI contract is `errorCount > 0` ⇒ fail the build.

## Generate — Viewer

The publication verb: governed meaning becomes dbt model property files with enforced contracts, one per model, colocated beside that model's own `.sql`. Four optional body fields:

| Field | Type | Meaning |
|---|---|---|
| `typeNames` | array of string | Restrict generation to the named models — matched against the dbt model name or the governed type's label. Empty or omitted means everything eligible. |
| `targetVersion` | string | The dialect to emit for. Omitted means modern: the `data_tests:` key. `"1.7"`, or any `1.x` below 1.8, emits the legacy `tests:` key and records a lossiness note saying dbt 1.8 and newer read `data_tests:`. An unparseable version stays modern and records a note. Enforced contracts require dbt 1.5 or newer either way. |
| `spaces` | array of string | Target space ids; empty means the project's main space. |
| `extra` | map of string to string | Open vendor options — for dbt, `layout` and `iriInDescription`. Both values are strings. |

`extra.layout` decides how the emitted files land in your repo:

| `layout` | Artifact names | Files returned |
|---|---|---|
| `model` *(default)* | `models/staging/stg_orders.yml` — beside the model's own `.sql` | one per eligible model |
| `folder` | `models/staging/_coremodels__models.yml` | one per model directory |
| `single` | `models/coremodels_contracts.yml` | one, repo-wide (back-compat) |

Colocation uses the model's `original_file_path`, captured at import into the dbt metadata mixin under the key `dbt.path`; a model imported before that capture, or a governed-first type with no dbt origin, falls back to `models/<model>.yml`. An unrecognized `layout` does not fail the call — it falls back to `model` and records a `SemanticNarrowing` at path `layout` naming the valid values.

`extra.iriInDescription` defaults to on. Governed `mapsTo` bindings — OLS-bound ontology terms, or any other standard — always ride out structurally under `meta.coremodels.maps_to`, on the model and on each column. With `iriInDescription` on they are *also* appended to the column `description` as `[standard: uri]`, because `description` is the only slot dbt's `persist_docs` carries into the warehouse column comment. Send `"false"` to suppress the description form; the structural form always rides.

```bash
curl -sS -X POST "https://coremodels.example.com/graph/integrations/dbt/generate/$PROJECT_ID" \
  -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \
  -d '{ "typeNames": ["orders","customers"], "targetVersion": "1.9", "extra": { "layout": "model" } }'
```

### Reading the response

```jsonc
{
  "success": true,
  "artifacts": [
    { "name": "models/marts/orders.yml",    "kind": "yaml", "content": "# Generated by CoreModels …" },
    { "name": "models/marts/customers.yml", "kind": "yaml", "content": "# Generated by CoreModels …" }
  ],
  "lossiness": [
    { "kind": "StructuralDrop", "path": "orders",
      "explanation": "This model already has properties in 'models/marts/_marts__models.yml'. Remove its block there, or dbt will report a duplicate patch for the generated 'models/marts/orders.yml'." },
    { "kind": "SemanticNarrowing", "path": "orders.parent_order_id",
      "explanation": "The reference's target field is not recorded and the target has no unique-tested column; relationships test omitted." }
  ],
  "errors": []
}
```

`artifacts[].name` is a **repo-relative path**, not a label — it is where the file belongs, and a writer script should create directories from it. Expect **many** artifacts: under the default layout a 400-model project returns 400 of them, each with `kind: "yaml"`. That is the point of per-model files — one repo-wide file is unusable past a few dozen models — but the response grows with the estate, so use `typeNames` when you want a subset.

`lossiness` is the list of things the generator refused to guess at. Each record is `{ kind, path, explanation }`, where `kind` is one of `StructuralDrop`, `SemanticNarrowing`, `TypeApproximation`, or `ConstraintRelaxation`, and `path` locates the model or `model.column` involved. Three are worth routing to a human: an ephemeral model skipped, a model with no columns skipped (dbt rejects an enforced contract with none), and the duplicate-patch record above. That last one is deliberate — dbt refuses two property blocks for one model, and the file already holding that model's `patch_path` may carry properties for models CoreModels does not govern, so the generator names the file whose block must be removed rather than overwriting a file it does not own. Decoding one `content` string shows what your PR will contain:

```yaml
# Generated by CoreModels — governed model contracts.
# Meaning changes belong in CoreModels; regenerate this file rather than editing it.
version: 2

models:
  - name: orders
    description: "One row per placed order."
    config:
      contract:
        enforced: true
      materialized: table
    columns:
      - name: order_id
        data_type: integer
        constraints:
          - type: not_null
        data_tests:
          - unique
          - not_null
      - name: status
        description: "Lifecycle status. [schema.org: https://schema.org/orderStatus]"
        data_type: "varchar(20)"
        data_tests:
          - accepted_values:
              values: ["placed", "shipped"]
        meta:
          coremodels:
            maps_to:
              - standard: "schema.org"
                uri: "https://schema.org/orderStatus"
            vocabulary: "orders status values"
```

`data_type` is on every column because an enforced contract requires it — the vendor-native type recorded at import, otherwise a warehouse-neutral fallback. Quoting is conservative: plain tokens stay bare, anything else is double-quoted, which is why `"varchar(20)"` is. `accepted_values` comes from a governed taxonomy, so one vocabulary generates the same list in every model using that field, and `meta.coremodels.vocabulary` names it. Governed references emit a `relationships` test with `to: ref('customers')` and `field: customer_id`; when the target field is unknown the test is omitted and a lossiness record says why, because a guessed field would generate a test that fails against a real warehouse.

### When Generate fails

Generate never emits an empty file. If no model survives selection, the call fails and says so:

```json
{ "success": false, "payload": null,
  "lossiness": [ { "kind": "StructuralDrop", "path": "int_order_items",
                   "explanation": "Ephemeral models cannot carry dbt contracts; skipped." } ],
  "errors": [ { "path": "generate", "message": "No eligible models found to generate contracts for." } ] }
```

The `lossiness` array survives the failure, so the skip records explain the empty selection instead of leaving you to guess. Two other shapes are worth recognizing: an import or audit body with no `artifacts` returns `errors: [{ "path": "artifacts", "message": "Body must include 'artifacts': { \"<name>\": \"<content>\" } (e.g. manifest for dbt)." }]`, and a mistyped vendor key returns `success: false` with `Unknown vendor '<v>'. Registered: <comma-joined keys>.` on every route.

## Status — Viewer

`GET graph/integrations/dbt/status/{projectId}` returns `{ success, vendor, imported, state, governedDatasets }`. `imported` says whether an import has ever run; `state` is the last-import record — `projectName`, `importedAt`, `toolVersion`, `artifactVersion`, `generatedAt`, `sourceFingerprint`, a human-readable `counts` summary, and `facts` carrying what the parser noticed, metric and exposure and semantic-model counts among them; `governedDatasets` is how many imported dbt objects currently resolve to governed Types. Comparing `sourceFingerprint` against a fresh manifest's is the cheapest way to know whether anything moved.

Beside these four, the generic integration surface carries `GET graph/integrations/vendors` for discovery, `reaudit`/`history`/`badge` for the audit trail, `POST graph/integrations/reconcile/{projectId}` (Admin) to link a dbt model and the warehouse table it materializes as one entity, the sync-plan routes, and the API-key surface at `POST v1/{projectId}/integrations/dbt/audit` that a CI gate calls. Note what none of them are: there is no schema diff between two points in time, no notification when someone edits meaning in CoreModels, and no UI for any of this — HTTP and MCP are the surfaces today. A worked walkthrough of these calls, from first manifest to first generated contract file, lives in the dbt quickstart that ships with the CoreModels integration docs.
