Get task-adapted open modelsfrom your model calls.

An abstract weave of orange, pink and blue pixels.

hello, middleware.

Enter

Middleware generates lightweight model updates from your calls and feedback. Built for fast refresh as your workload changes.

Join the waitlist (opens in a new tab)

$100 in free model credits to try it out.

Between your application
and its model calls.

Keep the tools around your application. Send its model calls through Middleware, then use your existing providers or open models from Mubit.

Your application keeps its frameworks, observability, evaluation and retrieval tools. Model calls pass through Middleware. Mubit provides a pool of open models connected to Middleware. Existing providers remain accessible directly or through your router or gateway.

Agent frameworks

LangGraphCrewAI

Observability

LangfuseLangSmith

Evaluation

PromptfooBraintrust

Context & retrieval

LlamaIndexQdrant

Your application

Prompts and tools

Model calls

Open models

Hosted by Mubit

Middleware

by Mubit

One base URL change.

Direct APIs

or your router / gateway
Your providers
  • OpenAI
  • Anthropic
  • Google + others
Illustrative overview.

Why Middleware comes first.

Middleware uses application memory to refresh task-specific model behavior. Placing it before the router keeps that process independent of provider selection. Your router still handles provider access, retries and fallbacks.

Mubit will host the open models and generate lightweight updates as your workflows change. The base models stay fixed. Use them alongside your existing providers, with the same application setup.

Models, context and feedbackBlue, orange and green circles change arrangement as the article scrolls.
Models, context and feedback.

Your work has
specific requirements.

Model APIs made it easier to add text generation, document extraction and other language tasks to an application. Developers could build a first version with a prompt and an API call.

Running those applications often requires repeated corrections. A document parser may misread the same field. A support tool may need the same wording changed before a reply can be sent.

Your application’s memory records those requirements, along with conventions, policies and examples of completed work. We’re building Middleware to use that changing record to refresh model behavior, while keeping the base model fixed.

Access to more models through an API

Teams can choose from hosted models and open-weight models, then compare their performance on the same workload. Open-weight models also allow developers to train and run versions suited to a particular task.

OpenRouter provides access to multiple providers through a common API. This reduces the integration work needed to compare models or switch providers.

Model selection still requires testing on your own inputs. A model that performs well on general benchmarks may need additional examples or training to handle your document formats, terminology or output requirements.

Context improves individual responses.

Context engineering involves selecting the instructions, documents, examples and tools a model receives for a request. These can supply current facts, explain required formats and let a model interact with an application.

For document extraction, context might include the expected fields and examples of correctly processed documents. It can improve a response without changing the model’s trained behavior.

Once the task is complete, you can record whether the output was accepted, what needed correction and whether the task succeeded. Reviewed examples can then be used to train a model for similar work.

Refresh behavior as your workflows change.

A specialized model needs to keep up with changes to your application’s conventions and policies. Our research studies how to generate lightweight model updates from application memory, while keeping the base model fixed.

A shared model component carries common behavior. A generator produces a small, application-specific update on top of it. A change detector monitors memory and triggers a refresh when enough has changed.

For example, when your escalation policy changes, the refresh process can update the model’s escalation behavior. Current facts and exact records still come from context and retrieval. Memory remains the source of record.

The research’s demonstrated path uses a brief fitting step before generating the update. Generating updates for a new application without that fitting step is still under development. The focus is frequent behavior updates without retraining the base model.

We’re bringing this refresh process into Middleware, so applications can use updated, task-specific models alongside their existing providers.