Most IT work is pointless.
I take on the part that isn't: hard software and data problems for organizations doing work that matters.
Something resembling a plan.
Every engagement starts with the best plan we can make — and the near-certainty that real data will punch holes in it. What stays fixed is how I work: in the open, in small pieces, with you close enough to steer.
Demos, not decks.
You judge working software running against your data, early enough that changing course is still cheap.
Dead ends included.
I write down what didn't work and why. Read the PoliLoom devlog to see what that looks like.
You keep the keys.
Code, documentation, and a team that understands what's running. Built to hand over, not to be needed.
What I do
Fifteen years of production work keeps circling back to the same set of capabilities:
LLM extraction & search
Structured facts from unstructured documents, and agents that do the research. Underneath both is hybrid search, reconciling records across sources and languages over millions of entities.
Interfaces for working with data
Review screens, research tools, search frontends: interfaces where people judge, correct and explore. Built for people whose work depends on the data being right.
The full stack
Database, pipelines, API, the interface people work with, and the servers it runs on—auth, backups, deploys and monitoring included. One person, end to end.
Stay up to date
New models, tools and techniques, evaluated against real problems and adopted when they earn it. Right now that means coding agents and LLM tooling; next year it will mean something else.
From roadmap to rollout
Find the hard part
Every project has one. We go there first, with real data.
Try the smallest useful thing
A prototype that does something real, within weeks.
Learn where it breaks
Real users, real edge cases, honest notes.
Make it boring
Tested, monitored, documented: software your team can run without me.
Current work
- Project
- PoliLoom, part of EveryPolitician: structuring politicians' data for investigators and the accountability sector.
- Client
- OpenSanctions
- Role
- Project lead, EveryPolitician (2025–present)
The problem
Assemble and verify structured politician data from Wikipedia/Wikidata and the wider web, across languages, ensuring provenance, correctness, and scale.
Solution highlights
Two-stage extraction pipeline: LLM extracts free-text positions → hybrid search maps to existing entities → LLM reconciles.
Fast hybrid search: Meilisearch with OpenAI embeddings for combined semantic and lexical entity matching
Source verification: web sources archived as MHTML via Playwright and reviewed through a FastAPI + Next.js confirmation UI
Impact
Clarity: From unstructured source documents to structured, linkable records.
Trust: Every extracted fact links back to a specific passage in an archived snapshot of the source.
Scale: Handles Wikidata-sized inputs through an incremental, parallelized pipeline.
LLM entity reconciliation actually works well, and with human-in-the-loop verification, it's both accurate and accountable. Read the devlog, or the Wikimedia Deutschland interview about the project. The Kolkhoz & Pravda projects ask the next question: can the same ideas work on any page — and how do you find the pages worth reading?
Writing & speaking
I think in public — essays on this site, talks at conferences and meetups.
Latest talks
- Who is running the world?
- Wikimania 2026, Paris — with Ada Homolova ·
- Text embeddings: navigating text in high dimensions
- Dataharvest 2026 — with Ada Homolova ·
- PoliLoom: Verification-First AI for Political Data in Wikidata
- WikiDataCon 2025 — with Brenna Maeve ·
- Finding connections
- Road to NODES 2025 (Neo4j) ·
- Finding connections: transform your document collections into a graph visualisation
- Dataharvest 2025 — with Lasse Edfast ·
About Johan
I am an autodidact software and data engineer with fifteen-plus years of experience, most of it for organizations doing public-interest work. I use LLMs to accelerate development, but never at the expense of clarity, reliability, or ethics.
I work remotely, Europe-focused but global clients welcome.
![]()
Selected experience
- OpenSanctions
- Project lead, EveryPolitician (2025–present)
- Follow the Money
- Full Stack Developer (2021–2025)
- Forest.host
- Founder (2017–2021)
Let's get in touch
- https://www.linkedin.com/in/johanschuijt/
- GitHub
- https://github.com/monneyboi/
- johan@resolve.works
- Phone
- +31 651 952 461
Frequently asked questions
What kinds of problems are you best at solving?
Data problems where information is scattered, unstructured, or trapped in formats that don't talk to each other. Think: extracting structured facts from thousands of documents, connecting data across systems, or building pipelines that turn messy inputs into something reliable and searchable.
I use LLMs where they genuinely help—extraction, matching, classification—but they're usually one piece of a larger system. If your problem is better solved with a spreadsheet or a well-written SQL query, I'll tell you that.
How involved does our team need to be?
More at the start, less over time. Early on I need access to the people who understand the problem—what's actually painful, what the data looks like, what "good enough" means. That might be a few hours in the first week or two.
During prototyping I'll share work frequently and need feedback. Once we're building for real, involvement drops to occasional check-ins and testing. By handover, the goal is that your team understands what's running and can operate it without me.
What does a typical project timeline look like?
It depends entirely on the problem. A small integration might take a few weeks; a complex data pipeline with verification workflows takes months and evolves as we learn what actually works.
Rather than give you made-up estimates, I'd point you to the PoliLoom devlog—it shows how a real project unfolded, including the dead ends and course corrections. That's more honest than a tidy timeline.
What I can promise: I ship early and often. You'll see working pieces within the first few weeks, not a big reveal after months of silence.
Who owns the code?
You do. Everything I build for you is yours—code, configurations, documentation. I prefer to build things that could be open-sourced if you wanted, and I'll actively suggest it when it makes sense. No vendor lock-in, no proprietary dependencies that tie you to me.
Do you also build the user interface, or just the backend?
Both. I design and build the full system—data pipelines, APIs, and the interface people actually use. A clear UI isn't optional; it's what makes the difference between a tool that gets used and one that gets abandoned.
What do you charge?
People hire me when the problem matters and the result has to hold up—if the deciding factor is price, I'm probably not the right hire.
What do you need from us to figure out if we're a good fit?
A conversation about the actual problem—not a polished pitch, just what's frustrating and why it matters. I work best with organizations doing something meaningful: journalism, accountability, public interest, open data, or businesses that genuinely care about doing good work rather than just scaling revenue.
If your goal is "add AI to make investors happy," we're probably not a match. If you're trying to solve a real problem and want to understand what you're building, let's talk.