At Toyoko we kept hearing about JEV, the API behind TypeSafe, and we wanted to see for ourselves whether it deserved the reputation. Rather than writing a toy script, we gave the test a real job to do. The problem we picked is a familiar one: a search on PubMed for our keywords returns hundreds of new papers every week, and most of them are noise. That test became PaperJev. It takes papers we have already written, or a summary of our research themes pasted into a form, pulls the last week of new biology papers from PubMed, and returns the ones worth reading, each with a relevance score attached.

The code is at github.com/ToyokoLabs/PaperJev, GPL-3.0. It runs from the command line or from a browser interface, and both drive the same three steps: download recent papers, summarize the themes of our own work, and filter the downloads against those themes. In this post we focus on the third step, because that is where JEV does the work.

PaperJEV GUI

The JEV idea

Most LLM APIs are organized around generating text. You write a prompt, the model writes prose back, and your code is left to scrape a number out of a paragraph, hoping the format survives the next model update. JEV starts somewhere else. You send one state, which is any JSON document you like, plus a set of typed questions about that state, and you get typed answers back.

JEV has three question primitives. A noul is a yes or no question whose answer comes back as a value between 0 and 1, so “probably” is representable and your code owns the cutoff. A score question rates the state from zero up to N against an ordered list of criteria you define. A choice question picks one label from criteria you define, which fits things like classifying an abstract as methods paper, review, or clinical study.

How PaperJev uses it

The filter reads two files: papers.json, the PubMed download, and summary.md, the themes summary. For every paper it sends this state:

{
  "article": "the paper's abstract",
  "summary": "the themes summary"
}

and one noul question: “Is this paper helpful or related to the subject described in the themes summary?” The answer comes back as a number. Anything above 0.5 is kept, and the number itself is stored as a relevance_score field, so the output can be sorted or thresholded differently later without re-running anything.

The loop is deliberately serial, one request per paper. The SDK handles the rough edges: it retries twice on rate limiting and server errors, honoring the Retry-After header when the server sends one. With no concurrency in our code there is nothing to coordinate and nothing to break when the API has a slow moment.

Running it from the command line

Put your NCBI credentials and TypeSafe API key in config.json, and your Ollama settings in ollama_config.json (this is used for the summarization step). Then:

uv run pubmed_downloader.py
uv run pdf_summarizer.py --dir /path/to/your/pdfs
uv run relevance_filter.py

The first command (main.py) downloads recent papers into papers.json. The second (pdf_summarizer.py) reads a directory of your own PDFs, summarizes them through Ollama, and writes summary.md. The third (relevance_filter.py) reads both files, sends one JEV request per paper, and writes the survivors to relevant_papers.json with their scores.

If you would rather not depend on the TypeSafe API at all, relevance_filter_LLM.py is a drop-in alternative that asks the same question using Ollama instead. Same two inputs, output file and cutoff value.

The browser version

All three steps also run in a web interface, which is what most people will actually use. There is a published Docker image, dnalinux/jevpapers:latest:

docker compose pull && docker compose up -d

then open localhost:8000. You upload your PDFs or paste an existing theme’s summary, set how many papers to pull from PubMed, and start. After downloading and filtering, the results page lists the relevant papers with their scores. The image is built for both amd64 and arm64.

Limits

One thing to be straight about: the PubMed search is broad. It queries “biology” over the last seven days, and none of your themes touch the query itself. The personalization happens entirely at the filter step, which is the part JEV is good at. A theme-driven search query would shrink the download, at the cost of trusting a keyword extraction to be as thorough as a broad pull plus a careful filter. We kept the broad search for now, and the scores in the output make it easy to judge whether that trade is right.

Since we wanted a real comparison, we also built the alternative filter mentioned above, relevance_filter_LLM.py, which asks the same question through a standard LLM, and ran both over the same download of about a thousand papers. The gap in time was hard to ignore. JEV finished in under five minutes. DeepSeek 4.1 flash took more than 15 minutes for the same job. Kimi K-2.6 needed more than three hours, and when we started a run with Kimi K3 we ran out of our allocated budget partway through and had to cancel it. Cost was the reason we watched these numbers so closely in the first place: a thousand JEV requests of this size come to about $0.08, while the LLM-based runs cost anywhere from $0.70 to $7 depending on the model.

The results were mostly the same set of papers, but the errors pointed in different directions. The LLM filters produced more false negatives, meaning relevant papers scored below the cutoff and were dropped. While JEV let some irrelevant papers through as false positives. For a literature watch this distinction matters, because with the LLMs there is a real chance of missing information, while with JEV the price of an error is just a few extra abstracts to skim.

Image credit: Biennale Architettura 2025 Book – Intelligens. Natural. Artificial. Creative

A first run with real papers

In Toyoko Bio, we explore emerging research areas and cutting-edge tools related to climate change. One of our areas is the environmental microbiome and probiotic buildings. To automate updates for our paper database, we conducted a first end-to-end test of the tool using real research papers. We downloaded all of them as PDFs into one directory (/home/toyoko/envpapers).

Before processing that directory, the first step was to pull the recent PubMed side of the experiment, using the downloader script with no parameters, which fetches the basic information for 1000 papers:

uv run pubmed_downloader.py

That produces papers.json, which we committed as a fixture at testdata/papers.json. This is a sample entry 


{
        "pmid": "42779000",
        "title": "γ-Aminobutyric acid-mediated regulation of osmolytes, nitrogen metabolism, and antioxidant defense confers dose-dependent drought tolerance in mungbean.",
        "journal": "Journal of the science of food and agriculture",
        "abstract": "Drought severely limits mungbean (Vigna radiata) growth by disrupting physiological and metabolic processes. Although γ-aminobutyric acid (GABA) enhances abiotic stress tolerance, its dose-dependent effects across different drought intensities in mungbean remain unclear. This study evaluated the effects of exogenous GABA (0, 0.5, 1, and 2 mmol L GABA enhances drought tolerance in mungbean through coordinated regulation of osmotic adjustment, antioxidant defense, and nitrogen metabolism, highlighting its potential as a priming agent for sustaining productivity under water-limited conditions. Future molecular studies are needed to elucidate the mechanisms through which GABA regulates drought tolerance in mungbean. © 2026 Society of Chemical Industry.",
        "pub_date": "2026-Sep"
    }

The second step ingests our PDF directory and produces the themes summary:

uv run pdf_summarizer.py --dir /home/toyoko/envpapers

Our output is at testdata/summary.md. Below is a preview of the contents of this file:

# Research Synthesis: Common Themes

# Common Themes Across the Research Papers

## Overarching Theme: The Microbiome of the Built Environment and Its Implications for Human Health

The papers collectively address the characterization, dynamics, and health consequences of microbial communities in human-constructed and occupied environments, spanning residential, institutional, healthcare, and urban atmospheric settings.

With summary.md and papers.json in place, we ran the relevance filter:

uv run relevance_filter.py

The result was testdata/relevant_papers.json, with 6 papers in it. We read those 6 manually and found 2 false positives, leaving 4 papers genuinely relevant out of the original 1000.

Here is a sample positive output:

{
        "pmid": "42776326",
        "title": "Occurrence of Stenotrophomonas carrying β-lactamase-encoding genes in a veterinary hospital environment.",
        "journal": "Veterinary research communications",
        "abstract": "Stenotrophomonas spp. are environmental bacteria increasingly recognized as opportunistic pathogens with multidrug resistance. This study characterized 17 environmental Stenotrophomonas isolates recovered from effluent samples and veterinary hospitalization units. The species identified included S. maltophilia, S. hibiscicola, S. muris, and S. forensis, with 13 distinct sequence types (STs) detected, including ST1402, which was shared between both environments. Antimicrobial susceptibility testing revealed a high resistance profile to ciprofloxacin, whereas resistance to tetracyclines, chloramphenicol, levofloxacin, and trimethoprim/sulfamethoxazole was less frequent. In accordance with the reported intrinsic resistance of Stenotrophomonas to a range of antibiotic classes, the strains were resistant to meropenem, ampicillin, cefoxitin, and amoxicillin/clavulanate. Whole-genome sequencing confirmed the universal presence of intrinsic β-lactamase genes blaL1 and blaL2, indicating the genetic potential for β-lactamases production, as well as the widespread distribution of RND-type efflux pumps, including components of the SmeABC and SmeDEF systems. In addition, all isolates were strong biofilm formers. Together, these findings indicate that environmental Stenotrophomonas spp. strains from veterinary settings harbor multidrug resistance determinants, efflux-mediated mechanisms, and biofilm-forming capacity, highlighting their potential role in the maintenance and dissemination of clinically relevant antimicrobial resistance.",
        "pub_date": "2026-Sep",
        "relevance_score": 0.7
    },

These are the two false positives:

{
        "pmid": "42779499",
        "title": "Two novel species of ",
        "journal": "Journal of helminthology",
        "abstract": "This study aimed to (1) assess the diversity of ",
        "pub_date": "2026-Sep",
        "relevance_score": 0.63
    }

{
        "pmid": "42772011",
        "title": "Radionuclides in the bottom sediments at the Komsomolets nuclear submarine wreck site (Norwegian Sea): fluxes and radioecology.",
        "journal": "Chemosphere",
        "abstract": "The vertical distribution of anthropogenic (",
        "pub_date": "2026-Sep",
        "relevance_score": 0.54
    }

Because their relevance scores exceeded 0.5, they were labeled as positives. This issue can be resolved by increasing the relevance score threshold to 0.65.