AI & customer experience

Where AI meets the contact center.

I spent twenty years running the operations that AI is now being sold into — global partner delivery, high-risk customer portfolios, enterprise platform migrations. I now build and evaluate AI systems on owned infrastructure.

That combination is the point. Understanding what a model does is one problem. Understanding the operating environment it lands in — and what has to be true before it works — is the other.

20+
Years enterprise operations
50 → 5,000
Global scale
50+
Specialized businesses
73
Controlled AI experiments
The bridge

Operator first, then the AI.

Contact-center and service-delivery organizations are where AI either earns its business case or quietly fails to. These four cases run in that order — what the systems do, then the operating environment they land in.

01 · Answer-engine measurement

Two controlled answer collections

exposed low recurrence and changed the methodology

Measuring how AI answer engines represent an organization

Condition

Organizations measure what they publish and what customers say back. Almost none measure what AI answer engines tell people who ask about them.

Role

Architect and sole operator — spec, schema decision, hand-curated panel, collection, and the companion value model.

Operating change

Sampling standard set empirically rather than by instinct: measured 9% recurrence at three samples, which killed single-shot collection and triggered a staged convergence experiment.

What it demonstrates

That he can evaluate model behavior, contain over-claiming, and build a business case without inventing a number — in his own words rather than a vendor's.

  • three gaps
  • two lanes
  • measurement ladder
View the evidence
Actions

Defined machine answers as a third collection surface rather than folding them into social listening. Scored three gaps: divergence from the organization's own account, from its ecosystem's, and assertions neither supports. Recorded two lanes per answer — trained belief versus live retrieval with citations. Derived prompts from the organization's own ecosystem, then curated by hand.

Outcome

324 answers collected in the first run, 400 in the second, across four engines. A companion value model — Decision Exposure, not cost — with a measurement ladder that must resolve to material mismatch before any figure touches money, and a published validation cascade naming which of its own stages remain unproven.

02 · AI system evaluation

73 experiments

produced deterministic scoring and identified a more reliable model architecture

Building a diagnostic that questions its own output

Condition

An LLM pipeline that scores organizational evidence is only useful if it is diagnostic rather than stochastic. Ask it the same question twice and get two answers, and it is a generator wearing an instrument's clothes.

Role

Designer, builder and operator. Published under his own name.

Operating change

Model and prompt changes stopped being judgment calls and became measured decisions with go/no-go gates — schema pass rate, unclassified rate, minimum finding counts.

What it demonstrates

That he can hold a technical conversation about model behavior, limitations and deployment architecture without being an engineer — and that when his own system produced plausible but inaccurate output, he stopped using it and rebuilt it rather than shipping around the problem.

  • deterministic scoring
  • model comparisons
  • reproducibility
View the evidence
Actions

Built extract, score, adversarial-skeptic and synthesize stages, where findings must survive challenge before entering the evidence ledger. Defined a correctness metric weighting reproducibility across repeated trials, evidence verification that checks whether cited identifiers actually exist, diagnostic value by finding type, and cross-run stability. Ran controlled model comparisons on a fixed corpus. Locked a production model standard with an explicit rationale, and a rule that scoring stays on the calibrated model even when extraction moves.

Outcome

Scoring determinism reached zero variance across identical runs. 73 experiments across four research tracks. A published head-to-head finding that the larger model did not win, and that prompt edits raising acceptance rates broke reproducibility — a tradeoff most teams never measure.

03 · Global delivery

50 → 5,000

partner agents in one year, including a new Cairo delivery site

Scaling a service operation from pilot to global delivery

Condition

A mobile product growing faster than the operation supporting it. Fifty agents needed to become five thousand within a year, across global partners, without customer experience degrading during the transition.

Role

Led the scaling and the new-market partner launch.

Operating change

A partner delivery model that held its quality bar while multiplying a hundredfold.

What it demonstrates

That he can build the delivery system behind a growth promise — the half most scaling stories skip.

  • work routing
  • quality ownership
  • partner operation
View the evidence
Actions

Treated it as a design problem rather than a hiring problem — establishing what had to be true before capacity was allowed to land: work routing, quality ownership, how a new site reaches standard rather than merely opening, and escalation paths across time zones. Stood up an entirely new partner operation in a market the organization had not operated in.

Outcome

50 to 5,000 agents in one year, including a new site in Cairo.

04 · Operating model

A unified decision model

across more than 50 specialized businesses

Installing decision rights across fifty-plus lines of business

Condition

Fifty-plus high-risk lines of business — regulatory, privacy, executive escalation, social, cybersecurity — each accumulated separately, each with its own governance and escalation path. No shared operating model.

Role

Designed and stood up the unified model.

Operating change

The question "who decides" acquired a documented answer in a portfolio where it had only ever had an informal one.

What it demonstrates

That he installs structure rather than becoming it — the difference between holding an organization together personally and building something that holds without him.

  • decision rights
  • risk tier
  • governance
View the evidence
Actions

Mapped the portfolio by risk tier, separating high-trust, high-consequence work from operational queues, then designed governance to match: explicit decision rights, defined escalation, named ownership per tier. Stood up a strategy function that had not existed, connecting frontline operations to enterprise decision-making.

Outcome

A unified operating model across 50+ lines of business, at roughly $200M operating budget, up to 1,300 internal employees and thousands of partner resources globally.

Enterprise platform adoption

A multi-year enterprise CX platform lifecycle, from the buyer's side

Owned the buyer's side of a multi-year enterprise CX platform adoption — the seat a platform vendor sells into. Migration completed against a dated 2019 target, forums transitioned on schedule in 2021, expansion into further operating groups the year after.

Canonical technical record

Experiments, benchmarks and published limits.

Greenbaum Labs holds the full record — methodology, results, and what each system does not do.

All selected work Experience LinkedIn Contact