Skip to main content

AppZen's ZenLM Plus Beats Frontier Models on Expense Audits

Image
AppZen's ZenLM Plus Beats Frontier Models on Expense Audits

San Jose, Calif. – October 01, 2026 -- AppZen has launched ZenLM Plus, a family of finance-specialized large language models that outperformed seven frontier AI models on expense audit accuracy while cutting inference costs, the company announced.

ZenLM Plus scores 97.4 F1 across targeted policy-category cases, beating the best frontier model by 8.7 points

AppZen tested ZenLM Plus against Opus 5, Sonnet 5, Gemini 3.1 Pro, Gemini 3.5 Flash, GPT-5.6 Sol, GPT-5.6 Terra, and GPT-5.6 Luna, giving each system identical expense data, supporting documents, and customer configurations. ZenLM Plus led five of six individual audit-control comparisons and all four consolidated policy groups, including a 10.8-point advantage in premium travel and upgrades, 4.2 points in electronic devices and gifts, and 2.5 points in policy and documentation exceptions. In targeted policy-category cases, ZenLM Plus reached an F1 of 97.4 versus 88.3 for GPT-5.6 Sol, the strongest frontier competitor.

Non-conforming receipt detection shows the widest performance gap, at 7.2 F1 points

ZenLM Plus scored 92.4 F1 in detecting non-conforming receipts, ahead of Opus 5 by 7.2 points. The model also led in cross-report duplicate detection, receipt itemization verification, and merchant category matching. Receipt verification was the single control where a frontier model edged ahead: ZenLM Plus posted 93.3 F1 against 94.1 for both Gemini 3.1 Pro and Sonnet 5.

Inference costs run up to 50 times lower than top-performing frontier models

ZenLM Plus delivered the lowest modeled inference cost among all systems tested. GPT-5.6 Luna, the cheapest frontier model evaluated, cost roughly twice as much per 1,000 audited expense lines, while Opus 5 cost approximately 50 times more. AppZen said the cost advantage makes specialized AI practical for high-volume audit workflows without sacrificing accuracy.

Mastermind Platform routes each finance task to the model best suited to perform it

AppZen's Mastermind Platform determines whether a given task should run on a ZenLM Plus model, deterministic logic, or a frontier model. "ZenLM Plus now outperforms the frontier models we tested on expense audit tasks while operating at a lower cost,

Published by
fairsonline_team
Industries
Company
Products
News Type