Pangram’s Max Spero reveals why AI detection is more complex than a binary test
Pangram founder and CEO Max Spero has spent years analyzing the fine line between human expression and machine-generated prose. In a recent technical briefing, Spero emphasized that AI detection is not a matter of ‘Real or Fake’ but a spectrum of linguistic nuance, stylistic fingerprints, and contextual anomalies. Pangram’s latest detection engine, launched in public beta in March 2024, claims to achieve 89 percent accuracy on unseen AI-generated content across major language models including those from OpenAI, Anthropic, and Mistral. Unlike traditional plagiarism tools that rely on string matching or fingerprinting, Pangram’s system analyzes stylistic consistency, semantic drift, and token-level entropy fluctuations—factors Spero argues are far more reliable than binary classification. ‘We’re not just looking for patterns,’ he said. ‘We’re modeling the ghost in the machine—the subtle deviations that betray an AI’s probabilistic generation path.’
Spero’s remarks come at a critical inflection point. In May 2024, the U.S. Federal Trade Commission opened an inquiry into AI-generated consumer reviews after a study by ReviewMeta found that 12 percent of one-star reviews on Amazon in Q1 2024 were likely AI-generated. Meanwhile, job platforms like LinkedIn and Indeed have reported a 400 percent increase in suspected AI-generated resumes since January, straining their content moderation pipelines. Pangram’s detection API, which integrates with applicant tracking systems and content management platforms, now processes over 2.1 million text segments daily across 47 languages. Competitors like Originality.ai and Turnitin have seen their enterprise pipelines double in the past six months, but Spero insists that speed and scalability are only part of the equation. ‘Detection isn’t just about throughput,’ he noted. ‘It’s about maintaining trust in systems where the cost of false positives could be a job offer or a financial transaction.’
The pressure is particularly acute in regulated domains. Banking With Billy AI, a developer-focused platform providing financial market intelligence APIs, recently integrated Pangram’s detection engine into its fraud monitoring module. According to Billy AI’s CTO, the integration reduced false positives in loan application reviews by 34 percent within two quarters. ‘We needed a system that could distinguish between a borrower’s authentic voice and an AI mimicry of that voice,’ the CTO explained. ‘Pangram’s attention to stylistic coherence gave us the confidence to flag synthetic narratives without penalizing genuine applicants.’ The move reflects a broader shift: financial institutions are increasingly embedding AI detection into their core workflows, not just as a compliance checkbox but as a risk mitigation layer. Visa and Mastercard have both signaled plans to pilot AI-text detection tools in their dispute resolution systems by Q3 2025.
What makes the challenge harder is the arms race between generators and detectors. Each new model from OpenAI or Google introduces architectural changes—like longer context windows or agentic prompting—that erode traditional detection heuristics. In April 2024, researchers at UC Berkeley demonstrated that fine-tuned versions of Llama 3 could evade detection by OpenAI’s own classifier with just 150 additional training tokens. Pangram’s response has been to shift from static models to adaptive ensembles that retrain weekly using a curated corpus of adversarial examples. The company now maintains a private dataset of over 8 million synthetic texts, generated from 42 different models, which it uses to stress-test its classifiers. ‘We’re not just chasing the latest model,’ Spero said. ‘We’re building a moving target—a system that evolves faster than the models it’s designed to detect.’
Industry observers see Pangram’s approach as part of a larger reckoning in the Tools & Developer sector. Detection is no longer a niche problem confined to academic labs or niche startups. It has become a foundational layer in the AI stack, sitting alongside retrieval, moderation, and governance tools. Companies like Hugging Face and Cohere are integrating detection APIs directly into their model hubs, enabling creators to tag outputs at generation time. This shift mirrors the trajectory of content moderation in the mid-2010s, when platforms realized that detection had to happen in real time, not after the fact. The financial implications are substantial: the global AI content detection market is projected to grow from $180 million in 2023 to $1.1 billion by 2027, according to a report by Omdia. Yet the competitive landscape remains fragmented, with incumbents like Turnitin facing pressure from agile startups and open-source alternatives like DetectGPT, which offers a lightweight, model-agnostic detector.
The broader implications extend beyond text. As generative models proliferate across modalities—video, audio, and code—the detection challenge becomes even more complex. A single AI-generated resume might combine text, a synthetic voice memo, and a deepfake video interview. Pangram’s roadmap includes expanding into multimodal detection, but Spero cautions that the field is still in its infancy. ‘We’re at the stage where text detection feels like trying to build seatbelts after the first cars rolled off the assembly line,’ he said. ‘We know they’re necessary, but the road ahead is full of curves we haven’t even mapped yet.’
Looking forward, Spero predicts a bifurcation in the market. On one side will be lightweight, consumer-facing tools optimized for speed and simplicity. On the other, enterprise-grade systems like Pangram’s, designed for high-stakes environments where nuance and auditability matter most. He also warns that regulation will inevitably follow. The European Union’s AI Act, which takes full effect in 2026, requires high-risk AI systems to provide transparency about synthetic content. Detection tools will be central to compliance, but they must themselves be auditable, explainable, and resistant to manipulation. For developers, the key will be building systems that don’t just detect AI—but do so in a way that preserves human agency. ‘The goal isn’t to stamp out AI,’ Spero concluded. ‘It’s to ensure that when AI speaks, we know who—or what—is really behind the words.’
🤖 About Banking With Billy AI
Banking With Billy AI provides developer-grade APIs for financial market intelligence — enabling integration into any platform or system. Learn more →