One flagship project, ten instruments, and a co-build program. Open any of them for the engineering underneath — the architecture, the workflow, and the honest answer to whether you need it.
The project · in development, components live
MBP Core — the AI-native open operational system
Not another tool on the pile — the layer that makes the pile think. MBP Core collects what your business already produces, resolves it into one living picture, and returns judgment-ready answers — while your people keep the final word. Built vertical-first, proven in transportation and logistics, designed for any operations-heavy business.
WHY
Every company already produces the signals it needs to run better — then loses them across inboxes, spreadsheets, and department walls. Core exists because the foundation IS the product: a structured store plus a semantic layer over your reality, compounding daily. And because closed platforms hold your operation hostage — Core is open: your systems remain the source of truth, every answer cites its evidence, and leaving with your capability intact is a designed feature, not a negotiation.
WHERE
Beside your existing stack, never as a rip-and-replace. Connectors read what you already run; a canonical relational store becomes the single source of truth; a semantic index (RAG over your corpus) sits on top. Your tools stay. Your data stays. The intelligence layer is what's new.
WHEN
After the groundwork conversation, before the tenth disconnected point-solution. If you're operations-heavy and your truth lives in more than three systems, you're the profile it was designed for.
HOW
Every request — asked or unasked — runs the same eight-stage lifecycle: Trigger → Resolve (classification, entity resolution, schema knowledge, semantic retrieval — the hard engineering lives here) → Gather → Assemble → Reason → Gate → Deliver → Log. The Log is the compounding asset: every human approval and correction is captured with reason codes and feeds back into stage five, so the system learns your judgment — a corpus no competitor can buy. Autonomy is graduated per agent per action (0–4: off, shadow, suggest, supervised, autonomous), every action carries an inverse operation, and money, customers, and compliance stay human-gated at every level. Forever.
IF
If a dashboard would solve it, buy a dashboard. Core earns its weight when the problem is the connections — when no single tool can see across your seams because the seams are between the tools.
IN PRACTICE Tuesday, 6:58am. Overnight: 240 emails, 31 documents, 4 sensor alerts, one customer complaint. By 7:00 every document is classified and extracted, the complaint is resolved against its order history, and the morning brief holds three items ranked by cost: a supplier price drift with the evidence attached, a certification window closing, a schedule conflict two weeks out. Two drafted responses wait at the gate. A human reads, corrects one, approves both — and the system just got smarter about next Tuesday.
Coordination layer · early access
MBP Room
A coordination layer on top of the systems you already run — because most operational failure isn't inside a team, it's between teams.
WHY
Work that crosses a team boundary loses its owner, its deadline, and its state. Rooms exist because the seam — not the department — is where operations leak money.
WHERE
On top of your existing stack, never instead of it. Each team keeps its tools; the Room owns what travels between them.
WHEN
When two or more teams routinely hand work to each other and "did you get that?" is a real sentence in your building.
HOW
Every cross-team transfer becomes a handoff object: a typed record with an owner, a state machine (created → accepted → done → verified), and an SLA clock. An event-driven routing layer assigns each object by rules first, model second; anything ambiguous or overdue escalates to a person. Command centers per team render only what that team owns — plus every clock currently running against them.
IF
If your operation is one team in one room, you don't need this — a shared board is enough. Rooms earn their complexity at the second seam.
IN PRACTICE Field crew completes a job at 4:52pm Friday. The completion photo becomes a handoff object owned by billing, with a 24h SLA. Monday 9am it hasn't been accepted — the Room escalates to the billing lead with the full chain attached. Nothing depended on anyone remembering.
Applied AI · active
Applied AI
Narrow, accountable intelligent systems aimed at the exact points where work and money slip — each one measured against a baseline taken before it existed.
WHY
Broad "AI transformation" fails at the rate of ~everything; a narrow system with a named KPI succeeds because it can be proven wrong cheaply. Accountability is the feature.
WHERE
Back-office seams: intake triage, exception handling, document reconciliation, follow-up chasing — wherever volume is high, rules are learnable, and errors are countable.
WHEN
After the groundwork exists (Layer 01/02 of the system). Never as the first move on a process nobody has mapped.
HOW
Each system is a pipeline: classification of the inbound signal, per-type extraction into structured fields, a policy step (deterministic rules first, model judgment second — in that order, always), then a drafted action that passes the human gate before execution. Every run is logged with inputs, confidence, and outcome; every system carries its KPI, its Week-0 baseline, and a stop condition. Autonomy is graduated: shadow → suggest → supervised.
IF
If a process changes weekly or depends on judgment you can't articulate, it's not ready for a machine — it's ready for a playbook first.
IN PRACTICE Inbound vendor invoices: classified on arrival, line items extracted, checked against history. 90% flow to the queue pre-matched; the 10% with anomalies — a duplicate, a price 40% off pattern — arrive flagged with the evidence attached. Six weeks in, the KPI reads against its baseline, not against a feeling.
Methodology · available
Playbooks
The Forge method written down — for teams who'd rather build it themselves. The groundwork, the seams, the gates, and the stop conditions everyone skips.
WHY
Because the method is the product. Tools change; the discipline of mapping before building, baselining before measuring, and gating before scaling doesn't.
WHERE
Inside your team, run by your people. A playbook is a transfer of method, not a dependency on us.
WHEN
When you have internal capability that wants to build, and what's missing is the sequence — not the enthusiasm.
HOW
Each playbook is a staged procedure with entry criteria, exit criteria, and artifacts: process maps with the handoffs made explicit, a data-readiness checklist, KPI + Week-0 baseline templates, a shadow-mode pilot protocol, and a gate review with three honest outcomes — scale it, fix it, or kill it and write down why. RACI included, because unowned steps are how playbooks die.
IF
If nobody on your team has bandwidth to run it, a playbook becomes shelf-ware — bring us in to run the first cycle instead, then take the second yourselves.
IN PRACTICE An ops manager runs the intake-triage playbook: maps the inbox reality in week one, baselines response times in week two, pilots classification in shadow for two weeks, then sits the gate review. The numbers clear the threshold — it scales. The next playbook, her team runs without opening the manual.
Advisory & research · standby
MBP AI Labs
The research arm where new techniques meet a live operation before anyone calls them a solution. If it hasn't survived real work, it isn't ready.
WHY
The industry ships demos as products. The Lab exists so nothing untested ever reaches an operation that depends on it — the failure happens where it's cheap.
WHERE
A sandboxed environment beside the live system: full read access, zero write access. The Lab can see everything and touch nothing.
WHEN
Continuously in the background; on demand when a new technique, model, or vendor claim needs a verdict grounded in your data instead of a benchmark.
HOW
Candidates run against historical replay (what would this have decided last quarter?) and then in shadow mode on live traffic, logged but inert. An evaluation harness scores them against the Week-0 baseline on accuracy, cost, latency, and failure shape — how it's wrong matters more than how often. The promotion procedure is binary and documented: graduate to a supervised pilot, or delete with a written reason. Nothing lingers half-alive.
IF
If a vendor won't let their tool run in shadow against your baseline before you commit — that's your answer about the vendor.
IN PRACTICE A new extraction model claims better accuracy on scanned documents. The Lab replays it across last quarter's corpus: 2% better on clean scans, 11% worse on phone photos — which is 60% of real intake. Verdict: delete, with the reason on file. Cost of the experiment: a weekend. Cost of learning that in production: a quarter.
Intake intelligence · field-proven
Inbox Sentinel
An always-on triage agent for the mailbox your operation actually runs on — because inboxes are where operations silently die. In the before-state we measured, requests sat unanswered for three days; the Sentinel's reaction time is about sixty seconds, around the clock.
WHY
A shared inbox is a queue with no SLA, no owner, and no memory. Requests that arrive at 23:40 wait for whoever opens the mailbox first — and the sender is already talking to your competitor. The Sentinel turns the queue into a system with a measured reaction time.
WHERE
Any high-volume shared mailbox: intake, sales, support, accounts payable, recruiting. Anywhere "who's watching the inbox?" is a real question at 2am.
WHEN
The day inbound volume outgrows the person reading it. Below ~20 messages a day it's overkill; above it, every unwatched hour is measurable leakage.
HOW
A polling loop (~60s cadence) reads every inbound message including attachments — per-type extractors parse PDFs, scans, and photos into structured fields. A classifier types each message against your request taxonomy; a drafting stage composes the reply from your history and pricing context; a routing card lands in the team channel with the full evidence chain. The draft waits at the gate — a human always sends. Every run is logged: message, classification, confidence, action taken.
IF
If your inbox is personal and low-volume, filters are enough. The Sentinel earns its keep when volume, off-hours arrivals, or attachment-heavy requests make human-only coverage a lie.
IN PRACTICE An offer lands at 23:40 with a rate sheet attached as a photo. By 23:41 it's classified, the numbers are extracted and priced against ninety days of history, and a draft sits in the morning queue with a recommendation and the comparables attached. The competitor's inbox opens at 8am.
Spend intelligence · field-proven
Leak Radar
Anomaly detection over the spend that arrives as documents — where duplicate billing and price drift hide at line-item level and compound in aggregate. Proven on a corpus of over a thousand real invoices.
WHY
Nobody reads line 14 of a 14-page invoice, which is exactly where the second billing of the same job lives. Individually invisible, collectively a percentage of your entire spend — and it never announces itself.
WHERE
Accounts payable, maintenance and repair spend, vendor services — any category where money leaves based on documents rather than structured system records.
WHEN
When spend arrives as PDFs, scans, and photos, and the person approving them is approving stacks, not lines.
HOW
Every document passes line-item extraction (not header-level OCR — each line becomes a structured record), then entity resolution ties lines to the actual vendor, asset, and job even across inconsistent naming. A statistical baseline per work-type forms from your own history; new lines score against it. Duplicates, price drift beyond tolerance, and pattern breaks surface as flags with the comparable history attached — a person confirms, the vendor answers, and the confirmation feeds the baseline.
IF
If you already run a mature ERP with disciplined three-way matching, the radar overlaps it — check your match coverage first. Most operations think they have it; the documents say otherwise.
IN PRACTICE Two invoices, five weeks apart, different reference numbers, same repair. Header-level review passed both. The radar caught them because line extraction plus entity resolution collapsed them onto one asset and one work-type — and the baseline said this job doesn't happen twice in five weeks.
Market intelligence · live in the field
Watchtower
A versioned radar on the world outside your walls — because operations optimize inward and get blindsided from outside. Runs unattended on a weekly cycle, with watchdogs guarding its own freshness.
WHY
Your pricing, capacity, and buy/sell timing all depend on conditions you don't control and mostly don't watch — until they've already moved. The Watchtower makes the outside world a scheduled input instead of a quarterly surprise.
WHERE
On top of the public data your market already publishes — price indices, demand indicators, cost inputs — rendered as one dashboard your team actually opens.
WHEN
Volatile input costs, seasonal demand, or any market where the spread between reacting this week and reacting next week is real money.
HOW
A scheduled pipeline pulls each source on its own cadence, normalizes units and geographies, and computes deltas against the prior issue. Output is an issue-numbered, dated artifact — versioned like a publication, so "as of when?" always has an answer. A separate read-only staleness watchdog verifies the live issue's age on its own schedule and alarms if the pipeline silently died — because the most dangerous dashboard is a stale one that looks current.
IF
If your market moves quarterly and your contracts are annual, this is a newsletter, not a tool — read one instead.
IN PRACTICE Monday 6am, the new issue lands: a regional demand spike, up sharply against last issue's baseline. Quotes go out that morning already adjusted — days before the competition's pricing catches up to what the data already said.
Cross-system audit · field-proven
Ops Sweep
A scheduled, full-spectrum audit agent that reads every platform you run — at the same time — because the dangerous problems live in the gaps between systems, where no single dashboard looks.
WHY
Every system you run reports on itself. Nothing reports on the space between them — where the calendar disagrees with the commitment, the billing disagrees with the delivery, and the deadline exists in a document nobody opened this month.
WHERE
Across email, chat, spreadsheets, trackers, and calendars simultaneously — wherever your operational truth is scattered.
WHEN
On a weekly schedule, and on demand before any leadership meeting — the sweep is the pre-read nobody had time to compile.
HOW
Parallel readers sweep every connected platform concurrently, each extracting its domain's facts. The cross-reference stage is the product: deadline vs. calendar, commitment vs. capacity, billing vs. delivery, promise vs. record. Findings are cost-ranked and compiled into one brief with structured, queryable frontmatter — so sweeps trend over time and "is this getting better?" has a data answer. Delivered to the owner, on schedule, unattended.
IF
If your whole operation lives in fewer than three systems, a good checklist beats this. The sweep pays at the fourth system, and compounds with every one after.
IN PRACTICE One sweep, one morning: a compliance deadline six days out that lived only in a document, a resource double-booked across two commitments on opposite coasts, and five figures of uncovered work sitting on a board nobody had reconciled. None of it visible in any single system — all of it in the gaps.
Judgment capture · in development
Decision Ledger
The approval inbox as a product. Every AI proposal a human corrects is training data — and most companies throw it away. The Ledger keeps it, codes it, and feeds it back until the system decides like you.
WHY
Working data is transient; decisions are permanent. The pattern of what your experienced people approve, correct, and reject — and why — is the most valuable dataset your company produces, and today it evaporates in the moment. Captured, it becomes a corpus no competitor can buy.
WHERE
Wherever the human gate lives: every system's approval step, every drafted action awaiting a yes. The Ledger is the gate's memory.
WHEN
From day one of any intelligent system — the corpus only compounds if it starts. Retrofitting judgment capture after a year of gating is a year of judgment lost.
HOW
Every proposal is logged with its confidence and rationale. Every human action — approve, correct, reject — is captured with a controlled-vocabulary reason code and the field-level deltas of what changed. The record is append-only: an audit trail by construction. On a training cadence, the corpus feeds back into the models — corrections become counterexamples, approvals become confirmation — and proposal quality climbs measurably against its own history.
IF
If nothing intelligent runs in your operation yet, there's nothing to log — start with the first system, ship the Ledger with it. It's the one tool here that should never be added later.
IN PRACTICE Two hundred gated decisions in, the drafts stop making the mistake the team corrected most — a delivery commitment that ignored a customer's receiving hours. Nobody retrained anything by hand. The correction pattern was in the Ledger, the Ledger was in the loop, and Tuesday's draft simply knew.
Predictive maintenance · in development, components proven
Fleet Maintenance Radar
Maintenance records are the most honest data an operation keeps — every invoice is a confession. The Radar reads what was and what is to tell you what's going to be — so you act before you react.
WHY
Reactive maintenance costs a multiple of planned — in money, downtime, and missed commitments — and the drift toward reactive is invisible month to month, because the story is scattered across invoices nobody reads line by line. The economics of a fleet live in that drift.
WHERE
Asset-heavy operations: vehicle fleets first — the Radar was born in transportation and logistics — and any equipment population with a repair history and a future.
WHEN
When repair spend arrives as documents, the planned-to-reactive ratio is a guess, and the same breakdown keeps surprising people who shouldn't be surprised.
HOW
Ingestion reads every repair invoice and work order with line-item extraction and resolves each line to its actual asset, vendor, and work-type. Per-asset baselines form from your own history: cost curves, service intervals, repeat-repair signatures. Then the WAS joins the IS — usage, reported faults, telemetry where it exists — and the model projects the will-be: assets trending toward failure, spend drifting from planned to reactive, warranties about to expire unused. Every projection lands as a flag with the evidence attached, and a human decides. Act, before react.
IF
If your fleet is under ten assets, a spreadsheet and a good mechanic beat it. The Radar earns its keep where no one person can hold every asset's history in their head.
IN PRACTICE One asset's records show the same repair three times in five weeks at rising cost — invisible across three invoices from two vendors, obvious once the lines resolve to a single asset. The unit comes in on your schedule, not on the roadside's.
Co-build program · open
Build With Us
Most AI is built for organizations, then handed over — and rejected like a transplant. Build With Us inverts it: your organization is actively involved in constructing its own operational ecosystem, with us embedded as the method — not the owner.
WHY
Adoption is where outside-built systems die: what a team didn't build, it doesn't trust; what it doesn't trust, it routes around. What a team builds itself, it runs, defends, and extends. Ownership from day one beats the best handover ever written.
WHERE
Inside your operation — your people, your systems, your data. We embed alongside the work, not above it. The build happens where the work happens.
WHEN
When you have people with the appetite to build and you want the capability in-house rather than rented. It pairs with any instrument on this page — the question isn't what gets built, it's who owns the muscle afterward.
HOW
A joint build pod: your operators with our engineers and method, running weekly build cycles inside live work — never in a lab. Every component follows the five passes, and every session builds dual competency: your people learn to construct intelligent systems while constructing them. Everything is documented into your own playbook as it's made, and each component graduates on the autonomy ladder — shadow, co-run, yours. The program's designed ending is the same as the creed's: you outgrow us.
IF
If nobody in your organization has bandwidth to participate, start with a built engagement instead — and switch to co-build the day you hire for it. Co-building without your time invested is just buying with extra meetings.
IN PRACTICE A coordinator who had never written an automation ships the intake triage with us in week three. By month two, they're specifying the next build — without us in the room. That's the product.