Data and artificial intelligence
Applied generative artificial intelligence
I design useful, measurable generative AI use cases: document search, summarisation, writing assistance and controlled answers. Every feature is evaluated on real cases before being exposed to users.
What it covers
Generative AI rarely fails because of the model. It fails because the question is vague, because the available documents do not contain the answer, or because nobody defined what a good answer looks like before building the interface.
So we start with the use case: which decision-maker, which decision, what expected gain, what error tolerance. Then comes the document material: source quality, usage rights, update frequency and chunking into usable passages. The model and infrastructure come last.
Evaluation before deployment
We build an evaluation set of around one hundred representative questions, with expected answers and legitimate sources. Every system change is measured against it: share of correct answers, answers without sources, and cases where the system admits it does not know.
- Use case framing and business success criteria.
- Corpus preparation with quality control and usage rights.
- Document retrieval with mandatory source citations.
- Versioned evaluation set replayed on every change.
- Logging of questions, answers and consulted sources.
Problems addressed
Trials never get past the demonstration stage. Users cannot verify the model's answers. Source documents are outdated or unreachable for the teams. Usage costs rise with no link to produced value. Confidentiality requirements are incompatible with an external service.
Expected benefits
Traceable answers grounded in identified documents. A clearly bounded scope of use, explained to users. Quality measured objectively before and after each change. Usage costs tracked per feature. Integration into your existing applications, not another isolated tool. A viable exit plan if results do not meet expectations.
Method and steps
- 1
Use case framing
Interviews with target users to define questions in scope, out-of-scope answers and the expected gain in time or decision quality.
- 2
Document preparation
Source inventory, usage rights verification, removal of outdated versions, chunking into passages and vector indexing.
- 3
Chain construction
Implementing retrieval, generation with citations, filtering of out-of-scope answers and interaction logging.
- 4
Evaluation and industrialisation
Building the reference set, measuring quality, fixing weak points, then integrating into the application and tracking usage costs.
Deliverables
Use case and success criteria framing document.
Cleaned, indexed document corpus with rights tracking.
Retrieval and generation chain with citations.
Versioned evaluation set and measurement report.
Dashboard for usage, quality and costs.
User guide and known system limitations.
Technologies used
- Python
- Spring AI
- PostgreSQL
- pgvector
- Redis
- Kubernetes
- LangChain
- API de modèles de langage
Frequently asked questions
Should we use an external or a self-hosted model?
It depends on volume, budget and confidentiality constraints. A self-hosted model prevents any data leaving your infrastructure but requires dedicated hardware; an external model is cheaper to operate but requires a contractual framework and filtering of the data you send.
How do you prevent invented answers?
By requiring citations of the passages used, by refusing to answer when retrieval returns nothing relevant, and by measuring the share of unsourced answers against a reference set on every change.
What gain can we expect?
For document use cases, the gain shows mainly in search time and completeness of answers, with wide variation depending on corpus quality. Exact figures must be measured on your scope during a pilot.
Related case studies
Internal document assistant with guardrails and evaluation
Augmented retrieval over a corpus of standards and technical notes
Demonstration sample - not a real client reference.
Study my use case
Describe the business need and the documents involved. Together we will check whether a four to six week pilot makes sense.