Platform

One integration replaces three vendors.

Plinth sits between your application and every model you might want to call. It picks one, checks the answer against your evals, and writes down exactly what happened. You integrate once.

Inference

The router is the product.

Send a request to plinth-route-v2 and we choose the model. The choice is yours to constrain: cheapest that clears your eval bar, fastest inside a latency budget, or a fixed model when you need a fixed model.

When a provider degrades, requests move. You find out from the trace instead of from your users.

Routes

/extract

active

Routing policy

cheapest above 0.94

route:

/extract

policy:

cheapest_above

threshold:

0.94

latency_budget_ms:

800

eligible:

- atlas-8b

- atlas-70b

- orbit-pro

failover:

true

MODEL

COST

P50

EVAL

ELIGIBLE

atlas-8b

$0.60/M

180 ms

0.95

yes

atlas-70b

$3.40/M

410 ms

0.97

yes

orbit-pro

$2.90/M

520 ms

0.96

yes

orbit-mini

$0.55/M

140 ms

0.86

below bar

Automatic failover. A provider timeout retries on the next-best model inside your latency budget. You pay only for the call that succeeded.

Pinning. Name a model explicitly and we route to it and nothing else.

Streaming. Token by token, with the same failover semantics as a blocking call.

Structured output. Pass a JSON schema. We validate before returning and retry a malformed response for free.

Batch. Submit up to 50,000 requests, get a file back. Half price, no latency guarantee.

Evals

The regression you didn’t ship.

An eval is a set of inputs, a set of expected properties, and a grader. Plinth runs yours on every prompt change, model change, and route change, then fails the build when the score drops past your threshold. The bad version never reaches production.

Evals

contracts.yaml

run #1,204

▲ Gate failed. Score dropped 0.05 from the previous run.

0.87

threshold 0.92 · 340 cases

CASE

GRADER

EXPECTED

SCORE

bol-0042

schema

valid

1.00

bol-0117

similarity

≥ 0.90

0.71

bol-0203

exact

MAERSK

1.00

bol-0288

rubric

≥ 0.85

0.63

bol-0341

schema

valid

1.00

The change that caused it: one word in the system prompt.

Graders. Exact match, JSON schema, semantic similarity, or a model-graded rubric you write yourself.

Regression gates. Set a floor per suite. A failing gate returns a non-zero exit code in CI.

Drift detection. Suites re-run nightly against pinned models, so a provider-side change surfaces within a day instead of a quarter.

Golden sets. Promote a real production trace into a test case in one click.

plinth evals run ./evals/contracts.yaml \
  --model plinth-route-v2 \
  --gate quality:0.92

Observability

Every call, written down.

A trace records the request, the model that answered, tokens in and out, latency, cost, and the eval score if one ran. Filter by user, route, deploy, or anything you attach as metadata. No instrumentation to write.

Usage

August 2026

month to date

Spend by route

$20,600

/extract

$14,200

/summarise

$4,100

/classify

$1,600

other

$700

By model

share of spend

atlas-8b

61%

3.1M calls

atlas-70b

29%

412K calls

orbit-mini

8%

890K calls

orbit-pro

2%

31K calls

Cost attribution. Spend per feature, per customer, per deploy, out of the box.

Latency percentiles. p50, p95, p99, split by model and by route.

Replay. Re-run any historical trace against a different model and diff the output side by side.

Export. Stream traces to your own warehouse. We are not trying to become your log store.

Architecture

What actually happens to a request.

Six steps, and the only one we add latency to is measured in single-digit milliseconds.

01

Ingest

Catalog, stock, reviews and policies flow in from Shopify, so AI reads your real data even where crawlers skip JavaScript.

02

Resolve

Variants and duplicates collapse into one clean product record, so AI matches the right product to the question.

03

Enrich

Missing facts like material, fit and use case are filled and approved. They are what AI uses to choose.

04

Serve

Your store is published as structured data AI can read: product schema, llms.txt, clean titles and descriptions.

05

Act

Agents check stock and build carts through permissioned actions you control.

06

Learn

Every prompt you win or lose shows which product facts to fix next.

01

Ingest

Catalog, stock, reviews and policies flow in from Shopify, so AI reads your real data even where crawlers skip JavaScript.

02

Resolve

Variants and duplicates collapse into one clean product record, so AI matches the right product to the question.

03

Enrich

Missing facts like material, fit and use case are filled and approved. They are what AI uses to choose.

04

Serve

Your store is published as structured data AI can read: product schema, llms.txt, clean titles and descriptions.

05

Act

Agents check stock and build carts through permissioned actions you control.

06

Learn

Every prompt you win or lose shows which product facts to fix next.

01

Ingest

Catalog, stock, reviews and policies flow in from Shopify, so AI reads your real data even where crawlers skip JavaScript.

02

Resolve

Variants and duplicates collapse into one clean product record, so AI matches the right product to the question.

03

Enrich

Missing facts like material, fit and use case are filled and approved. They are what AI uses to choose.

04

Serve

Your store is published as structured data AI can read: product schema, llms.txt, clean titles and descriptions.

05

Act

Agents check stock and build carts through permissioned actions you control.

06

Learn

Every prompt you win or lose shows which product facts to fix next.

Integrations

Change one base URL. Keep the rest of your code.

Official SDKs for Python, TypeScript, Go, and Ruby. An OpenAI-compatible endpoint if you would rather not change anything at all.

py

Python

ts

TypeScript

go

Go

rb

Ruby

{ }

REST

>_

CLI

gh

GitHub Actions

dk

Docker

tf

Terraform

wh

Webhooks

db

Warehouse export

ot

OpenTelemetry

1M tokens a month, free forever.

No card, no trial clock, no sales call. Move to a paid plan when the invoice makes sense, or don’t.