Engineering with evidence

Building AI systems
you can put to the test.

I'm Jieyu. I turn AI product ideas into working prototypes, with evidence to judge what to build next.

I hold an MSc in IT & Cognition from the University of Copenhagen. My work connects retrieval, evaluation, guardrails, and observability with human-centred product design.

I can work independently, taking ownership of scoping, implementation, and validation while keeping collaborators informed. I am willing to relocate to Amsterdam for an applied AI or AI product engineering role.

Selected projects

Working prototypes, evaluation feedback, and the decisions behind them.

Concept illustration of a lens selecting a document from an archive
RAG & agentsEngineering project

RAGOps Lens

A RAG evaluation platform that compares retrievers, gates uncertain answers, and makes quality, latency, and cost visible.

70 evaluation cases, confidence-gated fallback, and SQL analytics.

  • FastAPI
  • pgvector
  • Qdrant
  • Azure
GitHub Opportunity Miner project illustration showing its evidence-first workflow
RAG & agentsFull-stack project

GitHub Opportunity Miner

An agent that turns GitHub issues into traceable product opportunities, buyer hypotheses, and concrete validation plans.

Source-linked opportunity cards, explicit decisions, and reproducible mock mode.

  • LangGraph
  • FastAPI
  • Next.js
  • GraphQL
Poster for the Fund Facts Cross-Check narrated demonstration
EvaluationCLI prototype

Fund Facts Cross-Check

Two models extract fund facts. A separate verifier checks their evidence, because agreement alone does not make a claim correct.

27 fixed evaluation cases and mutation checks on a synthetic corpus.

  • Python
  • Structured outputs
  • Regression evals
CareMind demonstration poster with a mobile care-journal interface
Human-centred AIHackathon project

CareMind

An edge/cloud care-agent project that helps dementia family caregivers organise observations and prepare for follow-up appointments.

Mobile demos, structured care records, and local-device privacy workflows. Non-diagnostic support.

  • Edge AI
  • Gemma
  • Cloud Run
  • Mobile
Archived experiment: 45 generated cases
Model configurationPassed
Qwen2.5-0.5B2 / 45
SmolLM2-360M6 / 45
Qwen2.5-1.5B16 / 45
EvaluationValidated MVP

FaultLine

A reliability harness for tool-using agents. Reproducible failure cases are checked against executable assertions and reference solutions.

45 generated cases across three model configurations, with archived traces.

  • Python
  • Agent evaluation
  • Programmatic checks

Gaze-informed retrieval

  1. Eye gazeExpert attention
  2. Anatomical priorsInterpretable bias
  3. RetrievalHybrid + reranking
Human-centred AICourse research

GazeRAG

A research framework using radiologists' eye gaze as interpretable anatomical priors for hybrid retrieval and reranking.

Query expansion, gaze-aware reranking, and retrieval sensitivity analysis.

  • BM25
  • Dense retrieval
  • Eye tracking
From prototype to decision

These cases distinguish technical evaluation from user validation. They do not claim measured delivery times or customer adoption.

GitHub Opportunity Miner — make an idea testable

Delivered: A runnable FastAPI and Next.js prototype that turns GitHub issues into source-linked opportunity cards, buyer hypotheses, and validation plans. A deterministic mock mode makes the workflow easy to demonstrate.

Evaluation feedback: The documented badcase suite checks for weak evidence, duplicate opportunities, and unsupported willingness-to-pay claims. The README records a live run with 38 GitHub items and one card passing its validation rules; that is a pipeline check, not proof of demand.

Decision: Keep evidence quality separate from commercial validation. The product distinguishes build, validate, watch, and reject recommendations; willingness to pay remains a hypothesis until tested with real users.

Inspect the evaluation approach ↗

Fund Facts Cross-Check — test the failure, not just the happy path

Delivered: An AI-assisted CLI prototype comparing two models' claims against synthetic fund factsheets, with saved results and a narrated demonstration. I selected the problem scope and model pair; Codex drafted the implementation and ran the checks.

Evaluation feedback: The documented checks cover 43 unit tests and 27 fixed cases. Deliberately disabling unit normalisation or bypassing the evidence gate causes regression failures. I personally compared the saved live claims and quotations with the synthetic sources and reviewed the mutation results.

Decision: Treat model agreement and source support as separate checks. Keep the prototype limited to controlled-language synthetic documents, with HTTP serving and authentication deferred rather than presented as finished features.

Read the contribution and AI-use disclosure ↗

Research

How do we tell whether better retrieval actually helps an agent?

MSc thesis

From retrieval alignment to realised utility

I studied experience retrieval for an LLM agent in TextWorldExpress CookingWorld, separating offline retrieval alignment from the agent's terminal success.

The central question: when a retrieval metric improves, does the agent actually perform better?

Thesis graphical abstract: de-lexicalised retrieval saturated offline alignment, while the terminal-success difference from raw semantic retrieval was small and uncertain.
Graphical abstract from the submitted thesis. Full protocols, outcomes, and verification code are available in the repository.

Background

University of Copenhagen

MSc, IT & Cognition

A foundation for work at the intersection of machine learning, language, cognition, and human-centred AI.

Learning in public: LLM Zoomcamp
AI systems
RAG, LangGraph, tool calling, structured outputs, guardrails
Evaluation
Golden datasets, retrieval metrics, regression testing, failure analysis
Engineering
Python, FastAPI, PostgreSQL, Docker, Azure, React, Next.js
More research & prototypes

Let's build something useful.

Interested in reliable AI systems, thoughtful evaluation, or an applied AI role? I'd be happy to connect.

Connect on LinkedIn