Article Bite

pg-jev Puts Jev Inside PostgreSQL's WHERE Clause—With an API Call Behind It

pg-jev makes Jev's probabilistic judgments callable from PostgreSQL, but its convenience comes with batched API requests containing row data, installation privileges, and data-sharing decisions.

Primary source pg-jev ↗ Published 2026.10.03

A SQL query can already answer whether a ticket is open. Could it also decide whether the customer sounds ready to cancel? pg-jev brings TypeSafe’s Jev decision model into PostgreSQL as ordinary functions. Its October 3, 2026, v0.2.1 release also permits a Jev-compatible local server without an API key. This is a community extension, not a TypeSafe product; it turns a natural-language condition into a model judgment, not a new kind of SQL index.

pg-jev project wordmark beside a cartoon PostgreSQL elephant holding a small data table.
pg-jev's original project header, shown as identification—not a diagram of its SQL or API behavior. Image from the project repository; © 2026 Zachi, PostgreSQL License and notice. Open original image ↗

A predicate that calls a model

A project example has the shape SELECT * FROM tickets WHERE jev(tickets, 'the customer threatens to cancel');. The first argument is the row, not a string column. jev() returns a boolean after applying a threshold to Jev’s yes/no probability. Other functions expose the probability (jev_prob), choose from named options (jev_choice), or score an ordered set of levels (jev_score). You can combine them with normal SQL filters and sorting, but the prose condition is a semantic judgment—not a substitute for exact joins, date comparisons, or arithmetic.

Under the hood, pg-jev serializes rows and batches questions into calls to TypeSafe’s POST /v1/systemone contract. Its default batch holds 20 rows; the project says larger batches became less reliable at locating the intended row in its own tests. It reads ahead on base tables or views, keeps HTTPS connections around, and caches answers per backend session. Anonymous subquery rows cannot use that read-ahead path. A LIMIT can stop future work, but already in-flight requests may still finish; the default cache and total session memory are not capped by the prefetch-row setting.

The author’s 2,000-row example reports roughly 3.5 seconds, 100 requests, and $0.012 on a first pass, then about 50 milliseconds from a warm session cache. Those are project-reported measurements for one setup, not an independent benchmark or a promise that your table will behave the same way. Cache hits do not make a new condition—or another database session—free.

The trust boundary is in the database server

The extension’s control file requires plpython3u, PostgreSQL’s untrusted Python language. Installation needs a superuser and PostgreSQL 14–17, according to the project; managed services that withhold those capabilities cannot simply enable it. The SQL syntax is convenient, but the server process is now making model API requests on behalf of a query. With the default TypeSafe endpoint, row contents leave the database for that service. The local-compatible endpoint added in v0.2.1 can keep those requests on your network, but only if you configure and operate a compatible server.

The project exposes jev.max_rows_per_statement and jev.max_chars_per_statement as spend guards; both default to off. Apply cheap filters and column projection within the table or view supplied to jev() where possible. Its read-ahead implementation can send rows that an outer WHERE filter or LIMIT would later discard. For sensitive or large tables, test the actual rows sent, use the statement caps, and inspect jev_stats() rather than assuming a simple-looking query is cheap or private.

Before trying it on real data

  • Check whether you control a PostgreSQL 14–17 server with plpython3u and superuser access. Test in a disposable environment before installing an untrusted-language extension.
  • Decide which row fields may cross the API boundary, and whether the hosted or a Jev-compatible local endpoint meets your data rules.
  • Start with a small, labeled sample. Compare jev_prob() with your expected decisions and tune thresholds around uncertain cases rather than trusting a boolean by default.
  • Set row and character caps; test whether the supplied relation narrows the data actually sent. Measure request count, cache growth, latency, and cost in your own sessions.

Sources