LLM Gateway (Routing and Fallback)
Problem: Tokens are expensive, and the two kinds do not cost the same:
- Input tokens are cheaper, because they can be processed parallelly.
- Output tokens are costlier, because of sequential processing.
Big models have higher cost, therefore easy queries should be handled by smaller models.
The gateway sits between the application and the providers, and the request flows through it:
- Application → (request) → LLM Gateway → Model Providers
Responsibilities of the LLM Gateway:
- Keys, budgets, rate limits
- Routing, which is solved by machine learning
- Fallback resilience, which is solved by distributed systems
- Caching
- Logging and observability
Routing is an ML problem, because deciding the toughness of a query is a semantics issue.
There are 2 types, each with its own representative approach:
- Routing: check the query, analyze its toughness, and send it to the appropriate model. Cons: no feedback to know if the query was routed correctly. RouteLLM is the example of this.
- Cascade: send to the smallest model. If the output is no good, increase the model strength. Cons: wasted generations, high latency. FrugalGPT is the example of this.
RouteLLM: Shrink the scale of the problem down to 2: the router's output is either strong or weak.
More specifically, the router outputs a probability:
which reads as "the probability that the strong model would beat the weak model on this query."
Compare this probability to a threshold :
Using is intuitive, because it acts as a dial. If is high, the strong model is used sparingly, and only when the router is sure this needs to be handled by the strong model.
Data to train a router comes from ChatBot Arena, where each record is one query, 2 responses, and the user prefers one.
Evaluation of the router happens by PGR, not by raw score, because the raw score is dependent on the underlying models.
Eg: weak score = 6.0, strong score = 9.0, and router score = 8.4.
This means the router captured 80% of the advantage the strong model has over the weak one.
If the router always sends to strong, then PGR = 1 (router lazily wins by sending every request to strong model), therefore you have to add a cost. Now you need to find the best (cost, PGR) curve, which you do by sweeping from 0 to 1.
Problems with router training:
Label sparsity: there are not enough labels between 2 specific models. Fix: tiering.