HOLIDAY CASE STUDIES EFFICIENCY

Internal benchmark / October 2026

11× more efficient AI.

Same question. 90.9% fewer tokens. We measured how efficiently AI answers a customer question across live Outlook, Xero and CRM data: once through a traditional MCP workflow, and once through the Holiday Environment. The model and underlying records were the same.

Company
Holiday's Environment
Systems
Live Outlook, Xero and CRM data; job records in a test environment
AI model
GPT-5.6 Sol for both approaches
Compared
Traditional MCP workflow and the Holiday Environment

The same answer at a fraction of the processing.

In the expanded comparison, the traditional workflow used 410,977 tokens and eight agent-facing calls. Holiday's two-call result used 37,295 tokens. Both approaches returned every required data point; Holiday also returned useful linked records.

90.9%fewer total tokens410,977 versus 37,295
11×greater token efficiencyExpanded comparison
75%fewer agent-facing callsEight versus two
115%answer completenessAll 67 required points plus 10 linked records

One customer email. Four systems to make sense of it.

A customer sends an email. To respond properly, the AI needs to identify the customer in the CRM, check relevant correspondence, inspect active jobs and find related invoices in Xero. It also needs to flag where the records disagree.

With separate tools, the AI has to reconstruct the relationship between the sender, contact, job and invoice as it moves between systems. Each call adds material for the model to process.

Same AI. Same data. Two architectures.

  1. Traditional MCPThe AI navigated inbox, CRM, accounting and job records through separate tools.
  2. Holiday EnvironmentThe AI queried connected business context through Holiday's permission-controlled environment.
  3. Expanded comparisonThe traditional workflow was extended from five to eight calls. Because it sought the same customer information, the original two-call Holiday result was reused as the comparison baseline.

What the measurements show.

Holiday used fewer tokens and fewer agent-facing calls in both comparisons. The expanded traditional workflow processed more information to reach the same required answer.

Customer-context retrieval results. The Holiday value in Test 2 is the reused Test 1 result.
Measure Traditional MCP Holiday Difference
Test 1 · tool calls 5 2 60% fewer
Test 1 · total tokens 268,625 37,295 86.1% fewer · 7.2×
Test 2 · tool calls 8 2* 75% fewer
Test 2 · total tokens 410,977 37,295* 90.9% fewer · 11×
Answer completeness 100% 115% 10 extra linked records

* Holiday's two-call, 37,295-token result from Test 1 was reused as the baseline for the expanded comparison.

More complete, with linked context.

A complete answer required 67 data points. The traditional workflow returned all 67, scoring 100%. Holiday returned all 67 plus 10 relevant linked records, including correspondence and jobs related to the invoices. Its score was 114.9%, rounded to 115%.

Scores above 100% mean the answer included relevant linked context in addition to everything required; they do not mean that required facts were counted twice.

How the comparison was made.

Scope
Holiday's own business: live Outlook, Xero and CRM records, plus job records in a test environment. Customer details are not reproduced.
Controlled
The same customer question, GPT-5.6 Sol model and underlying records were used for both architectures.
Measured
Agent-facing calls, total tokens processed and answer completeness against 67 required data points.
Source
Holiday Environment Efficiency Study, internal benchmark, October 2026.

Cached input tokens remain part of total token usage but may be billed at a reduced rate by the model provider. This study compares token volume, not a measured invoice cost or response-time result.

The test covered one customer-context task and a traditional tool catalogue of about 60 APIs. Other tasks, integrations and model configurations may produce different results. Larger-deployment savings are a hypothesis, not a measured outcome of this study.

Download the full study (PDF)

Why this matters for a business.

  • Less processing. Fewer tokens can reduce AI usage charges. Actual charges depend on the provider's rates and any cached-token discounts; Holiday passes provider charges through with no markup.
  • Fuller context. In this task, Holiday returned the required answer and linked records while the AI made two agent-facing calls.
  • Governed access. Holiday's environment applies permissions to business reads and records AI activity for review.

Read the complete method and results.

The PDF includes both comparisons, token details and the answer-completeness method.

Internal benchmark run on Holiday's own business, October 2026. Customer details from live systems are not reproduced. Results will vary by task and integration.