Easital product
StepVideo
Built and run by Easital
More than one model provider in a single pipeline. Language and speech models process the transcripts, step context and audio derived from each recording.
LLM development
Easital Technologies Ltd. is an LLM development company. We integrate large language models into software, fine-tune them when a base model falls short, and deploy them through hosted APIs or on servers you control. The work includes the parts that decide whether an LLM feature holds up: evaluation, control of the output and cost per request.
LLM development is the engineering of software that depends on a large language model: selecting the model, shaping what goes into it and what comes out, adapting it to your domain and running it at a known cost.
It rarely means creating a language model from the beginning, which takes data and computing power that few companies need to spend. It means working on four layers. The model: which one, and where it runs. The context: the instructions, examples and data sent with each request. The output contract: the format the software expects back and what happens when the model breaks it. And operations: latency, spend, monitoring and version changes.
Most complaints that “the model is not good enough” turn out to be problems in the context or the output contract. We find out which layer is at fault before recommending a different model or a fine-tune, because a new model is the most expensive way to fix a prompt.
Eight services, taken together as one project or separately when part of the system already exists.
Hosted and open-weight models run against your own tasks and compared on quality, latency and cost per request. The result is a recommendation per task, not one model for everything.
A service between your application and the model that holds versioned prompts, assembles context, handles retries and rate limits, streams responses and can fall back to a second provider.
Instructions and examples written and tested like code, and context assembled so that the model sees what it needs and nothing more.
Training-data preparation, fine-tuning runs and a comparison against the base model on a held-out test set. Easital has done model training and fine-tuning.
Typed function definitions and schema-validated responses, so the model’s output can drive software directly. This is the foundation of an AI agent.
Open-weight models served on your own hardware or private cloud, for data that cannot be sent to a public API. See private LLM deployment.
A test set for each task, automatic scoring where the answer can be checked by rule, human review where it cannot, and tracing of live requests.
Routing of easy requests to smaller models, caching of repeated context and limits on tokens per request. Offered on its own as LLM cost optimization.
An LLM project runs in seven steps. Fine-tuning comes fourth, after the cheaper options have been measured.
Real inputs with the outputs you expect, and a scoring method for each: exact match for extraction, a written rubric for free text, human review for tone.
A general hosted model with a plain prompt, scored on the test. We record quality, latency and tokens per request, so every later step has a number to beat.
Clearer instructions, a few worked examples, and retrieval of the facts the task depends on. These changes are cheap and reversible.
We fine-tune when the test still fails in a consistent way, or when a smaller model has to match a larger one for reasons of cost, speed or data location. The numbers from steps two and three decide.
A schema for every response the software consumes, validators, a maximum length, and a defined behavior when the model refuses, breaks the format or does not know.
Requests are routed by difficulty, repeated context is cached, token limits are set and bulk work is batched. We report cost per request before and after.
The model version is pinned, and the test set is run again before any upgrade, because a new version can change behavior. Live requests are traced, and alerts fire on error rate, latency and spend.
Three live systems depend on language models every day: two that Easital owns and runs, and one built for a client.
Easital product
Built and run by Easital
More than one model provider in a single pipeline. Language and speech models process the transcripts, step context and audio derived from each recording.
Easital product
Built and run by Easital
An LLM that acts on code. The AI chat and agentic code editor in Manob.ai change a project from a plain-language request. A multi-step agent task uses more of the model than a simple prompt, and Easital did extensive work on the token cost of this product.
Client project
Built by Easital for a client in the United States
Model choice exposed per agent. The Calldone site describes building agents with custom prompts, models, transcribers and voices, and its privacy policy describes a language model that reads a sample of an agent’s recent transcripts and writes a short brief.
Fine-tuning changes how a model behaves. RAG changes what a model knows at the moment it answers. Use RAG when the problem is missing or changing knowledge, and fine-tuning when the problem is format, tone or a narrow skill.
Retrieval-augmented generation (RAG) looks up relevant passages in your content and sends them to the model with the question. Fine-tuning continues the training of a model on your own examples. Prompt engineering, the third option, only rewrites the request.
| Prompt engineering | RAG | Fine-tuning | |
|---|---|---|---|
| What it changes | The instructions and examples in the request | The facts supplied with the request | The weights of the model |
| Use it when | The base model can do the task once it is told how | Answers depend on your documents or on data that changes | You need a consistent format, tone or narrow skill, or a smaller model |
| Updating knowledge | Edit the prompt | Re-index the source, with no retraining | Train again on new data |
| Can cite its sources | No | Yes, the retrieved passages | No |
| Work before launch | Low | Medium: ingestion, index, evaluation | High: training data, training runs, evaluation |
| Cost per request | Grows with prompt length | Higher, since retrieved text adds tokens | Can be lower: shorter prompts, smaller model |
| Typical failure | Inconsistent on unusual inputs | The wrong passage is retrieved | Knowledge goes stale and is costly to correct |
The two can be combined: retrieval supplies current facts, and a fine-tuned model uses them in the format you require. Our order of work is prompt first, retrieval second, fine-tuning third. See RAG development.
Our LLM fine-tuning services start with the question of whether to fine-tune at all. It is worth the effort when the task is stable, you have good examples of it, and a tested prompt on a base model still misses the target.
We are not tied to one model vendor. Models are chosen per project by test, so everything here is given by category.
All major commercial model providers. Several are tested on your task, and the integration sits behind one interface so that the provider can be changed.
Easital takes on LLM work in three forms, from a complete project to a review of what you already have.
One LLM feature or one fine-tune taken through the seven steps above, delivered with its test set, its prompts and a deployment you can operate.
Best for: a task you can describe with examples.
Engineers from Easital who work in your repository and your process on integration, evaluation or fine-tuning. See hire AI developers.
Best for: teams that are building on LLMs already and need more capacity.
We test your feature against a set of real cases, locate the failing layer (model, context, output contract or operations) and deliver a written list of fixes in priority order.
Best for: features with inconsistent output or a bill that keeps rising.
An LLM development company builds and runs software that depends on large language models. It selects models by test, integrates them into applications, designs the prompts and the output format, fine-tunes a model when that is justified, and deploys it through a hosted API or on private servers with monitoring of quality and cost.
Fine-tuning changes the model’s weights by training it on your examples, which changes how it behaves. RAG leaves the model unchanged and supplies relevant passages from your content with each request, which changes what it knows. RAG suits knowledge that changes and answers that must cite sources. Fine-tuning suits a fixed format, tone or narrow skill.
A service layer between your application and the model: versioned prompts, context assembly, structured outputs with validation, retries and rate-limit handling, streaming, logging of every request and a fallback provider. Your application calls this layer, so the model behind it can be changed without changes to the application.
Pairs of input and the output you want, checked by someone who knows the domain. Existing records often provide them: resolved tickets, edited drafts, labeled documents. The amount depends on the task, so we start with a small set, measure the gain and add data only while the score keeps improving.
Yes. Open-weight models can be served on your hardware or in your private cloud, and a fine-tuned model can be deployed the same way. Easital has set up local and on-premise models. See private LLM deployment.
Nothing changes until it has been tested. We pin the model version in production and run the full test set against the new version first. If scores hold, the upgrade goes out in stages. If they drop, the prompts are adjusted or the older version stays in place for as long as the provider supports it.
Send a description of the task and, if you can, a few example inputs and outputs. We reply by email with questions and a proposed first step.
Easital is an AI and SaaS engineering company that takes AI software to production, and runs AI products of its own. Founded in 2019.