Most IT work is pointless.

I take on the part that isn't: hard software and data problems for organizations doing work that matters.

Contact me See my work

Something resembling a plan.

Every engagement starts with the best plan we can make — and the near-certainty that real data will punch holes in it. What stays fixed is how I work: in the open, in small pieces, with you close enough to steer.

  • Demos, not decks.

    You judge working software running against your data, early enough that changing course is still cheap.

  • Dead ends included.

    I write down what didn't work and why. Read the PoliLoom devlog to see what that looks like.

  • You keep the keys.

    Code, documentation, and a team that understands what's running. Built to hand over, not to be needed.

What I do

Fifteen years of production work keeps circling back to the same set of capabilities:

  • LLM extraction & search

    Structured facts from unstructured documents, and agents that do the research. Underneath both is hybrid search, reconciling records across sources and languages over millions of entities.

  • Interfaces for working with data

    Review screens, research tools, search frontends: interfaces where people judge, correct and explore. Built for people whose work depends on the data being right.

  • The full stack

    Database, pipelines, API, the interface people work with, and the servers it runs on—auth, backups, deploys and monitoring included. One person, end to end.

  • Stay up to date

    New models, tools and techniques, evaluated against real problems and adopted when they earn it. Right now that means coding agents and LLM tooling; next year it will mean something else.

From roadmap to rollout

  1. Find the hard part

    Every project has one. We go there first, with real data.

  2. Try the smallest useful thing

    A prototype that does something real, within weeks.

  3. Learn where it breaks

    Real users, real edge cases, honest notes.

  4. Make it boring

    Tested, monitored, documented: software your team can run without me.

Current work

Project
PoliLoom, part of EveryPolitician: structuring politicians' data for investigators and the accountability sector.
Client
OpenSanctions
Role
Project lead, EveryPolitician (2025–present)
  • The problem

    Assemble and verify structured politician data from Wikipedia/Wikidata and the wider web, across languages, ensuring provenance, correctness, and scale.

  • Solution highlights

    Two-stage extraction pipeline: LLM extracts free-text positions → hybrid search maps to existing entities → LLM reconciles.

    Fast hybrid search: Meilisearch with OpenAI embeddings for combined semantic and lexical entity matching

    Source verification: web sources archived as MHTML via Playwright and reviewed through a FastAPI + Next.js confirmation UI

  • Impact

    Clarity: From unstructured source documents to structured, linkable records.

    Trust: Every extracted fact links back to a specific passage in an archived snapshot of the source.

    Scale: Handles Wikidata-sized inputs through an incremental, parallelized pipeline.

LLM entity reconciliation actually works well, and with human-in-the-loop verification, it's both accurate and accountable. Read the devlog, or the Wikimedia Deutschland interview about the project. The Kolkhoz & Pravda projects ask the next question: can the same ideas work on any page — and how do you find the pages worth reading?

Writing & speaking

I think in public — essays on this site, talks at conferences and meetups.

Latest articles

The tool is not the author
AI agents are human too

All articles

Latest talks

Who is running the world?
Wikimania 2026, Paris — with Ada Homolova ·
Text embeddings: navigating text in high dimensions
Dataharvest 2026 — with Ada Homolova ·
PoliLoom: Verification-First AI for Political Data in Wikidata
WikiDataCon 2025 — with Brenna Maeve ·
Finding connections
Road to NODES 2025 (Neo4j) ·
Finding connections: transform your document collections into a graph visualisation
Dataharvest 2025 — with Lasse Edfast ·

About Johan

I am an autodidact software and data engineer with fifteen-plus years of experience, most of it for organizations doing public-interest work. I use LLMs to accelerate development, but never at the expense of clarity, reliability, or ethics.

I work remotely, Europe-focused but global clients welcome.

Profile shot of Johan Schuijt

Selected experience

OpenSanctions
Project lead, EveryPolitician (2025–present)
Follow the Money
Full Stack Developer (2021–2025)
Forest.host
Founder (2017–2021)

Let's get in touch

LinkedIn
https://www.linkedin.com/in/johanschuijt/
GitHub
https://github.com/monneyboi/
Email
johan@resolve.works
Phone
+31 651 952 461

Frequently asked questions

What kinds of problems are you best at solving?

Data problems where information is scattered, unstructured, or trapped in formats that don't talk to each other. Think: extracting structured facts from thousands of documents, connecting data across systems, or building pipelines that turn messy inputs into something reliable and searchable.

I use LLMs where they genuinely help—extraction, matching, classification—but they're usually one piece of a larger system. If your problem is better solved with a spreadsheet or a well-written SQL query, I'll tell you that.

How involved does our team need to be?

More at the start, less over time. Early on I need access to the people who understand the problem—what's actually painful, what the data looks like, what "good enough" means. That might be a few hours in the first week or two.

During prototyping I'll share work frequently and need feedback. Once we're building for real, involvement drops to occasional check-ins and testing. By handover, the goal is that your team understands what's running and can operate it without me.

What does a typical project timeline look like?

It depends entirely on the problem. A small integration might take a few weeks; a complex data pipeline with verification workflows takes months and evolves as we learn what actually works.

Rather than give you made-up estimates, I'd point you to the PoliLoom devlog—it shows how a real project unfolded, including the dead ends and course corrections. That's more honest than a tidy timeline.

What I can promise: I ship early and often. You'll see working pieces within the first few weeks, not a big reveal after months of silence.

Who owns the code?

You do. Everything I build for you is yours—code, configurations, documentation. I prefer to build things that could be open-sourced if you wanted, and I'll actively suggest it when it makes sense. No vendor lock-in, no proprietary dependencies that tie you to me.

Do you also build the user interface, or just the backend?

Both. I design and build the full system—data pipelines, APIs, and the interface people actually use. A clear UI isn't optional; it's what makes the difference between a tool that gets used and one that gets abandoned.

What do you charge?

People hire me when the problem matters and the result has to hold up—if the deciding factor is price, I'm probably not the right hire.

What do you need from us to figure out if we're a good fit?

A conversation about the actual problem—not a polished pitch, just what's frustrating and why it matters. I work best with organizations doing something meaningful: journalism, accountability, public interest, open data, or businesses that genuinely care about doing good work rather than just scaling revenue.

If your goal is "add AI to make investors happy," we're probably not a match. If you're trying to solve a real problem and want to understand what you're building, let's talk.