A language model is an engine. Not a vehicle.
640×
Fewer model calls
Score behavioural segments, not individuals. Same coverage, a fraction of the spend.
19
Pipeline node types
Composable steps assembled into a workflow — no bespoke build per use case.
100%
Traceable outputs
Every result logged with its inputs, the rules applied, and its rationale.
0
Releases to retune
Brand voice, business logic and guardrails are human-readable text, not code.
Point a raw model at a million customer records and it will produce a million fluent paragraphs. It does not know which products exist in your catalogue, what commercial constraints apply, or why it chose what it chose.
Fluency was never the hard part. Governance is. Being right, being on-brand, defensible, and measurably better next cycle — at a volume no human can review.
Roughly 5% of the platform is the model call. The other 95% is what makes the output usable.
Under the hood
Nine things you don’t get out of the box.
Not prompt engineering. Purpose-built machinery — retrieval mathematics, classical statistical ML, computer vision and a durable orchestration layer.
Segmentation
Behaviour-driven dynamic segments
Records group by what they actually do, then re-form as behaviour shifts — no hand-maintained buckets going stale.
Retrieval
MMR-reranked example selection
Candidates are embedded, score-filtered, then reranked so the model sees examples that are diverse and on-target.
Business logic
Commercial constraint engine
Encode the rules that make an output profitable rather than merely plausible — category, margin and strategy all enforced.
Governance
Guardrails & output validation
Brand voice, legal limits and commercial rules are enforced as transparent, human-readable rules before anything ships.
Measurement
Rubric-based evaluation
Versioned scoring metrics prove quality moved and tune against the metric that matters, not a vague read-better feeling.
Statistical ML
Benchmarking & topic modelling
Twelve specialised benchmarkers plus custom topic clustering for reading themes, trends, sentiment and fatigue across a whole corpus.
Creative
Automated visual composition
Computer-vision logo detection, performance-informed placement and a rules-based composition strategy registry for finished, on-brand creative.
Insight
Emerging & fatiguing signals
Fourteen analysers surface what’s gaining traction and what’s burning out — feedback that feeds the next cycle, not a dashboard nobody opens.
Scale
Durable batch orchestration
Fault-tolerant workflow nodes, embedding caches, credential rotation and model failover for million-record runs.
The production line
What happens between your data and a shipped output.
01
Ingest & standardise
Source data is normalised through channel-specific adapters into a single governed schema.
02
Segment
Records cluster by behaviour so thousands of segments can stand in for millions of individuals.
03
Retrieve context
Relevant exemplars, knowledge passages and reference data are embedded, filtered and reranked before the model sees them.
04
Generate under constraint
The model runs inside a frame of brand, legal and commercial rules, producing ranked results and fallbacks.
05
Validate & reject
Every output is checked against catalogue validity, template structure and rule compliance before it can ship.
06
Score & tune
Results are graded against versioned rubrics and real-world performance so accuracy compounds run over run.
The arithmetic
Hold the model constant and change only the approach: one call per record, versus scoring behavioural segments through the pipeline.
Underlying model
One call / record
Through TxtGen
Reduction
GPT-5 mini
~$1,800
~$30
60×
Gemini 2.5 Pro
~$9,000
~$130
69×
Claude Sonnet 5
~$18,000
~$240
75×
Or build it yourself
Segmenting first is the obvious optimisation — and it’s available to anyone. The catch is that by the time you’ve built the segmentation, retrieval, guardrails and evaluation harness, you have rebuilt most of TxtGen.
Behavioural segmentation & clustering
6–10 weeks
Embedding, retrieval & reranking
4–6 weeks
Guardrails, validators & rule engine
4–8 weeks
Evaluation harness & rubrics
6–8 weeks
Durable batch orchestration at scale
8-12 weeks
Traceability, logging & observability
4-6 weeks
Channel adapters & delivery integration
3-5 weeks
Where it runs
The expensive part is the platform, so the marginal cost of a second or third use case is small. Adding a programme is usually another agent and another step — not another system.
Offer & incentive personalisation
Individual-level offer selection across a full customer base, weighted for incremental margin.
Creative production at scale
Copy and finished visual assets generated per variant, per channel, per market — on-brand and validated.
Lifecycle & campaign messaging
Governed message generation across email, push and in-app, with rules marketing edits directly.
Product & catalogue content
Descriptions, attributes and localisations produced under catalogue-validity and brand rules.
Market & competitive insight
Theme, trend, sentiment and fatigue detection across large corpora without leaning on the model alone.
Decision support
Operational recommendations that are traceable to the data, rules and evidence behind them.
The real exposure
An expensive output costs you the token price once. A wrong output — off-brand, off-catalogue, non-compliant, or quietly discounting something that would have sold at full price — costs you margin, trust, and sometimes a regulator’s attention. Governance is not overhead. It is the product.
