skip to content
The Weighted Average

Enterprise AI & Work

Pangram 4 Makes AI Detection Ten Times Pricier

Pangram 4 raises scanning cost from $50 to $500 per million words, making buyer-corpus testing and appeals part of the purchase.

A stack of papers on a wooden table
A stack of papers on a wooden table. Photograph by 2H Media

Pangram 4 lists a million-word developer-API scan at $500, up from $50 under Pangram 3—a 10× unit-price increase before Pangram’s advertised bulk discount. Its new $9 million round equals 18 billion words at that list price; schools, publishers, recruiters, and platforms should shadow-test before buying the score without the appeals machinery that makes it safe.

Better vendor claims, a much larger bill

Pangram’s technical report claims 99.66% AI-text detection and a 0.0041% human false-positive rate. Its model card describes the evaluation boundaries and is the more useful procurement document: aggregate performance cannot guarantee behavior on a buyer’s languages, document lengths, edited text, or unusual human styles. Both figures remain vendor claims until independently reproduced.

The price change is concrete. Pangram’s pricing page lists Pangram 4 at $0.05 per 100 words and lists Pangram 3 at $0.05 per 1,000 words. At one million words, divide by the respective billing units: 10,000 blocks times $0.05 equals $500 for Pangram 4; 1,000 blocks times $0.05 equals $50 for Pangram 3. Those are standard developer-API list prices; the page advertises a 20% bulk discount, but applying the same discount to both versions preserves the 10× relative increase.

Pangram 4 makes the list-price scan ten times costlier

Developer API cost to scan 1 million words, before bulk discounts

Pangram 4Pangram 3$0$100$200$300$400$500$50$500
Pangram 4Pangram 3$0$200$400$50$500
Pangram Labs pricing page · Jul 2026

That is the budget line buyers should model before changing production traffic. A university screening 100 million words annually moves from $5,000 to $50,000 in raw API charges before integrations, human review, support, or appeals. The example is arithmetic, not a claim about a particular institution’s volume; procurement should substitute its own corpus.

Pangram’s launch post also introduces an image detector as a research preview. VentureBeat reports the text-and-image expansion alongside a $9 million round. Buyers should keep preview image results out of enforcement until the model has a public, independently evaluated operating record.

The false-positive arithmetic matters more than the headline accuracy. Apply the technical report’s vendor-claimed 0.0041% rate to 10 million genuinely human documents: 10,000,000 × 0.000041 equals 410 false flags. That number does not predict a buyer’s actual rate, because document distributions differ. It demonstrates why “roughly one in 24,000” can still create hundreds of cases at platform scale.

TechCrunch reports the launch alongside a $9 million funding round and found human-written sentences marked as AI-assisted. Divide that funding by the official $0.05-per-100-word list rate: it equals 18 billion words of Pangram 4 scanning before bulk discounts. That two-source figure is not capacity the company promised to provide; it shows the scale of capital entering a market where each uncertain result can still demand human adjudication.

The contrast with Microsoft’s capex-heavy AI cloud is useful: falling model-generation costs do not guarantee falling governance costs. Detection vendors spend compute and data on keeping pace with new generators, then customers spend labor adjudicating uncertain scores. The AI economy can make content cheap while making trust expensive.

Shadow mode first, enforcement only after appeals

Who should switch? Existing Pangram customers with a real detection need should test version 4 before changing endpoints, but they should not silently replace a production classifier. New buyers should compare probabilistic detection with provenance systems, process controls, and targeted human review. An institution that cannot tolerate false accusations should use the score as one investigative signal, never dispositive proof.

The cost model needs four lines: API spend, integration, reviewer time, and appeal handling. At the standard developer-API list rate, one million words cost $500 according to the official pricing page before its advertised bulk discount; a false flag can cost far more if it triggers faculty time, an employment dispute, content removal, or reputational damage. The earlier analysis of efficiency rationing applies here in reverse: optimize the expensive judgment loop, not merely model throughput.

A proper shadow test samples the buyer’s actual traffic across language, length, genre, accessibility style, non-native writing, paraphrase, and human-edited model output. Reviewers should be blinded to the detector score. The team should report false-positive and false-negative rates with confidence intervals, then establish a threshold tied to the harm of each error. A global benchmark cannot choose that threshold for a school or publisher.

Appeals must exist before enforcement. Preserve the submitted artifact, model version, threshold, score, and reviewer rationale. Give the affected person a way to provide drafts, revision history, sources, or other provenance. Prohibit automated punishment from one detector result. These controls are not bureaucratic decoration; they are part of the product’s effective price.

What could break the thesis? Independent testing could show Pangram’s very low false-positive claim holds across difficult subgroups, making the higher price economical relative to review saved. Conversely, new generators and editing techniques could erode detection quickly, buyer-corpus false positives could exceed tolerance, or provenance adoption could make probabilistic detection less valuable. Distribution drift is the central technical risk.

Evidence that changes the verdict is measurable: an independent evaluation on the organization’s corpus, subgroup results, stable performance over model updates, a documented retirement/migration path, and an observed cost per correctly resolved case below alternatives. If those tests pass, Pangram 4 may justify the tenfold rate. If they do not, the buyer should keep version 4 in shadow mode and spend the budget on provenance and review.

The sharp conclusion is not that AI detection never works. It is that a vendor-reported 0.0041% error rate and a $500-per-million-word price still leave the buyer responsible for every consequential decision. Buy the signal only with the machinery to challenge it.

Sources