RAG development

RAG development services: answers grounded in your own data

Easital Technologies Ltd. builds retrieval-augmented generation (RAG) systems: software that finds the relevant passages in your documents and records, hands them to a language model and returns an answer with its sources. Our RAG development services cover ingestion, retrieval, answer generation, evaluation and the upkeep that keeps answers correct as your content changes.

What retrieval-augmented generation is

Retrieval-augmented generation is a way of answering a question with a language model using facts looked up at the moment of the question, instead of relying on what the model memorized during training.

A RAG system works in two phases. Ahead of time, your content is split into passages and indexed so that it can be searched by meaning and by keyword. At question time, the system searches that index, selects the most relevant passages and gives them to the model with an instruction to answer only from them and to cite them.

This solves two things a language model cannot do by itself: know your private content, and know what changed yesterday. Updating what the system knows means updating the index. The model is not retrained.

RAG is the right tool when answers live in documents, help articles, tickets, contracts or product data, and a reader needs to verify the source. It is the wrong tool when the answer requires a calculation over a whole dataset, which is a database query, or when the problem is tone and format, which calls for fine-tuning.

Where RAG systems fail

Most RAG failures are retrieval failures, not model failures: the right passage never reached the model, or an outdated one did. The five problems below tend to appear once real users arrive.

Common RAG failures and how we build against them
FailureWhat the user seesWhat we do about it
Stale indexAn answer quotes a policy that has since been replaced.Scheduled and event-driven re-indexing, removal of deleted documents, and a last-updated date on every source.
Bad chunkingThe answer misses a condition because a passage was cut in half or a table was flattened.Splitting along the document’s structure, keeping headings with each passage, and testing chunk sizes against real questions.
Weak retrievalThe right document exists and is not found, often for exact terms such as part numbers.Hybrid search that combines meaning with keywords, metadata filters and re-ranking of the results.
No evaluationNobody can say whether the latest change made answers better or worse.A question set with known sources and answers. Retrieval and answer quality are scored separately on every change.
Ungrounded answersThe model fills a gap with a plausible invention.Answers restricted to the retrieved text, citations on every claim and an explicit “not found” reply.

What we build

A RAG system is a chain of six parts, and its answers are only as good as the weakest. We build all six, or repair the failing ones in a system you already have.

  • Document ingestion

    Pipelines that read PDFs, web pages, help articles, tickets, transcripts and database records, then clean them, remove duplicates and keep the metadata retrieval needs.

  • Chunking and indexing

    Passages cut along headings, sections and table rows instead of at a fixed length, then stored in a vector index and a keyword index side by side.

  • Hybrid retrieval and re-ranking

    Search by meaning and by exact term, filtered by metadata such as product, date or language, with a re-ranking step that puts the strongest passages first.

  • Cited answer generation

    A model that answers only from what was retrieved, links each claim to its source and says so when the content does not contain the answer.

  • Permission-aware retrieval

    Each user’s search covers only the documents that user may open. The filter runs inside the search, so restricted text is never sent to the model.

  • Sync and freshness

    Connectors that pick up new, changed and deleted content from the source systems, so the index does not drift away from the truth.

How a RAG build runs

A RAG build runs in eight steps. Retrieval is built and measured by itself before a language model writes a single answer, because a model cannot repair a search that returned the wrong passage.

  1. Inventory the sources

    Which content the system should answer from, its format, its owner, how often it changes and who is allowed to read it.

  2. Write the question set

    Real questions from users or support logs, each paired with the document that answers it and a reference answer. It includes questions the content cannot answer, so that “not found” is tested too.

  3. Build ingestion

    Parse, clean and chunk the content, then read a sample of the chunks by eye. Problems with tables, scans and repeated headers are cheapest to fix here.

  4. Build and measure retrieval

    For each test question we check whether the right passage is among the top results. Chunking, hybrid search, filters and re-ranking are tuned until it is.

  5. Add answer generation

    A grounded prompt, citations and the “not found” behavior. Answers are scored for correctness and for whether each claim is supported by the cited passage.

  6. Add guardrails and permissions

    Access filtering per user, handling of personal data, and treatment of retrieved text as data and never as instructions, so that a document cannot redirect the model.

  7. Control cost and latency

    The number and length of passages sent per question, caching for frequent questions and a smaller model for rewriting queries. Each change is checked against the question set.

  8. Release and monitor

    Every question is logged with the passages retrieved and the answer given. Unanswered questions are reported as content gaps, and production failures join the question set.

Proof from our own work

Three live systems answer or search from their own content. Easital owns and runs two of them and built the third for a client.

  • Easital product

    StepVideo

    Built and run by Easital

    Answers from a user’s own content, with sources. The StepVideo site describes an “Ask your library” feature: ask a question in chat and get the answer from your own guides, with every source linked.

  • Client project

    Calldone

    Built by Easital for a client in the United States

    Retrieval during a live phone call. The Calldone site lists knowledge base queries among the tools a voice agent can use, alongside transfers, SMS and webhooks.

  • Easital product

    Manob.ai

    Built and run by Easital

    AI search inside a marketplace. The Manob.ai site describes a smart search across marketplace products and services, assisted by AI.

See all of our work

Agentic RAG: when one search is not enough

Agentic RAG puts a language model in charge of retrieval: it decides what to search for, reads the results, and searches again or calls another tool until it has enough to answer.

Standard RAG runs one search per question. That breaks down on questions that need facts from several places, such as a comparison of the refund terms in two versions of a contract, and on questions where the second lookup depends on the result of the first.

Each extra search is another model call, so an agentic answer costs more, takes longer and is harder to test. We build standard RAG first and add the agentic loop only for the question types it cannot answer. The loop is an agent, with the step and spend limits described under AI agent development.

Standard RAG and agentic RAG compared
Standard RAGAgentic RAG
Searches per questionOneAs many as the model decides, within a limit
SourcesUsually one indexSeveral indexes, databases and APIs, each exposed as a tool
FitsDirect questions answered in one placeMulti-part questions, comparisons and follow-up lookups
Cost and latencyLow and predictableHigher and variable

Custom RAG development or a ready-made tool

A ready-made “chat with your documents” tool is enough for a small, clean set of files used by one team. Custom RAG development services are worth paying for when one of the conditions below applies.

  • The content lives in several systems and has to be kept in sync with them.
  • Different users are allowed to see different documents.
  • The documents contain tables, scans or a structure that generic splitting destroys.
  • The answers reach customers, so their accuracy has to be measured.
  • The documents may not be sent to a public service, which calls for a private model.

Technology we use for RAG

A RAG stack is chosen to suit the content and hosting rules of each project, so every part is given by category.

Language models
  • Commercial model providers
  • Open-weight models on private servers

All major commercial model providers and open-weight models are options. The choice follows your hosting rules first, then answer quality and cost on your question set.

Retrieval
  • Embedding models
  • Vector search
  • Keyword search
  • Re-ranking models
Ingestion
  • Document parsers
  • Text recognition for scans
  • Scheduled and event-driven sync
Evaluation and monitoring
  • Question sets
  • Retrieval hit-rate scoring
  • Answer grading
  • Request tracing

Ways to work with us

RAG work with Easital starts in one of three ways, depending on whether you are building, extending or repairing.

  • Scoped build

    One RAG system over a defined set of sources, taken through the eight steps above and delivered with its question set.

    Best for: a knowledge base, document archive or product catalog with a clear group of users.

  • Dedicated engineers

    Engineers from Easital who join your team to build or extend retrieval inside your own product. See hire AI developers.

    Best for: product teams adding search or question answering to existing software.

  • Audit of an existing RAG system

    We run your system against a question set, separate retrieval failures from generation failures, and deliver a written list of fixes in priority order.

    Best for: systems that answered well in a pilot and are now losing the trust of their users.

RAG development: questions and answers

What is retrieval-augmented generation in plain terms?

It is an open-book exam for a language model. Before the model answers, the system searches your content for the passages that relate to the question and hands them over. The model writes its answer from those passages and cites them, so the answer reflects your current content and can be checked.

What do RAG development services include?

Ingestion of your content, chunking and indexing, hybrid retrieval with re-ranking, answer generation with citations, access control, an evaluation set, monitoring and the sync that keeps the index current.

What is agentic RAG?

Agentic RAG lets the model run retrieval as a series of steps: it plans what to look up, searches, reads the results and searches again or uses another tool before it answers. It handles multi-part questions that a single search misses, at a higher cost and latency per answer.

Should we use RAG or fine-tune a model?

Use RAG when the model lacks knowledge, especially knowledge that changes or must be cited. Fine-tune when the model has the knowledge and gets the format, tone or a narrow skill wrong. The two can be combined. A comparison table is on the LLM development page.

Why does our RAG system give wrong answers?

The usual causes are a stale index, passages cut in the wrong places, search that misses exact terms, and no evaluation set, which hides the other three. Check, for a sample of failed questions, whether the right passage was retrieved. If it was not, the fault is in retrieval and a different model will not help.

Can RAG run without sending our documents to a public AI service?

Yes. The index, the embedding model and the language model can all run on your own servers or in your private cloud. Easital has set up local and on-premise models. See private LLM deployment.

Tell us what your users need answered, and from which content

Describe the sources, who asks the questions and what a wrong answer would cost. We reply by email with questions and a proposed first step.

Easital is an AI and SaaS engineering company that takes AI software to production, and runs AI products of its own. Founded in 2019.