
RAGOps Lens
A RAG evaluation platform that compares retrievers, gates uncertain answers, and makes quality, latency, and cost visible.
70 evaluation cases, confidence-gated fallback, and SQL analytics.
- FastAPI
- pgvector
- Qdrant
- Azure
Engineering with evidence
I'm Jieyu. I turn AI product ideas into working prototypes, with evidence to judge what to build next.
I hold an MSc in IT & Cognition from the University of Copenhagen. My work connects retrieval, evaluation, guardrails, and observability with human-centred product design.
I can work independently, taking ownership of scoping, implementation, and validation while keeping collaborators informed. I am willing to relocate to Amsterdam for an applied AI or AI product engineering role.
Working prototypes, evaluation feedback, and the decisions behind them.

A RAG evaluation platform that compares retrievers, gates uncertain answers, and makes quality, latency, and cost visible.
70 evaluation cases, confidence-gated fallback, and SQL analytics.

An agent that turns GitHub issues into traceable product opportunities, buyer hypotheses, and concrete validation plans.
Source-linked opportunity cards, explicit decisions, and reproducible mock mode.

Two models extract fund facts. A separate verifier checks their evidence, because agreement alone does not make a claim correct.
27 fixed evaluation cases and mutation checks on a synthetic corpus.

An edge/cloud care-agent project that helps dementia family caregivers organise observations and prepare for follow-up appointments.
Mobile demos, structured care records, and local-device privacy workflows. Non-diagnostic support.
| Model configuration | Passed |
|---|---|
| Qwen2.5-0.5B | 2 / 45 |
| SmolLM2-360M | 6 / 45 |
| Qwen2.5-1.5B | 16 / 45 |
A reliability harness for tool-using agents. Reproducible failure cases are checked against executable assertions and reference solutions.
45 generated cases across three model configurations, with archived traces.
Gaze-informed retrieval
A research framework using radiologists' eye gaze as interpretable anatomical priors for hybrid retrieval and reranking.
Query expansion, gaze-aware reranking, and retrieval sensitivity analysis.
These cases distinguish technical evaluation from user validation. They do not claim measured delivery times or customer adoption.
Delivered: A runnable FastAPI and Next.js prototype that turns GitHub issues into source-linked opportunity cards, buyer hypotheses, and validation plans. A deterministic mock mode makes the workflow easy to demonstrate.
Evaluation feedback: The documented badcase suite checks for weak evidence, duplicate opportunities, and unsupported willingness-to-pay claims. The README records a live run with 38 GitHub items and one card passing its validation rules; that is a pipeline check, not proof of demand.
Decision: Keep evidence quality separate from commercial validation. The product distinguishes build, validate, watch, and reject recommendations; willingness to pay remains a hypothesis until tested with real users.
Inspect the evaluation approach ↗
Delivered: An AI-assisted CLI prototype comparing two models' claims against synthetic fund factsheets, with saved results and a narrated demonstration. I selected the problem scope and model pair; Codex drafted the implementation and ran the checks.
Evaluation feedback: The documented checks cover 43 unit tests and 27 fixed cases. Deliberately disabling unit normalisation or bypassing the evidence gate causes regression failures. I personally compared the saved live claims and quotations with the synthetic sources and reviewed the mutation results.
Decision: Treat model agreement and source support as separate checks. Keep the prototype limited to controlled-language synthetic documents, with HTTP serving and authentication deferred rather than presented as finished features.
How do we tell whether better retrieval actually helps an agent?
I studied experience retrieval for an LLM agent in TextWorldExpress CookingWorld, separating offline retrieval alignment from the agent's terminal success.
The central question: when a retrieval metric improves, does the agent actually perform better?
MSc, IT & Cognition
A foundation for work at the intersection of machine learning, language, cognition, and human-centred AI.
Learning in public: LLM ZoomcampInterested in reliable AI systems, thoughtful evaluation, or an applied AI role? I'd be happy to connect.
Connect on LinkedIn