Intermittent context truncation / distillation
With the way that I understand Venice to work, the context of an inference is sent to a given GPU inference provider, which returns a response. On subsequent queries within the same “chat”, the whole context is sent to another inference provider, and so on.
If the initial prompt contains, say, a medical test results PDF which I ask the model to extract the results from, and that document contains my other personal information, then is it safe to say that the PDF contents are sent to the next inference provider for each request?
It would be useful to be able to click a button and have all context above that point hidden from future requests. Additionally, another mode for that button could be “distillation”, where the model is prompted to produce a scrubbed summary of the above context, such that subsequent requests will only include the distilled context.
The purpose of this being to limit the distribution of identifying information between difference inference providers.
Log in to comment and vote
Comments2
Silver River
Feb 4, 2025
I might be misunderstanding but I think this could be useful for an even broader set of scenarios. This would be great for interactive storytelling RP as well. Sometimes I have a question to ask and being able to hide it from the llm to save context for the story but still have the answer present would be useful. More so the summary would really help with the limited context and the need to restart conversations regularly to continue the story due to limited context window, general token limits and token management.
Black Woods
Feb 4, 2025
Agreed that this is broadly useful as a way of compensating for limited context if enabled for the whole chat instead of selectively as described above
The idea would be to separate what is presented to the user from the (more concise) context that accumulates during the chat. Basically, you’d be using two models—one for chat and one (presumably smaller and cheaper) for context compression