Please add prompt caching to Jev-style models (emschwartz.me)

🤖 AI Summary
A developer working on the personalized content feed Scour has proposed the addition of prompt caching to Jev-style "System One" models to enhance efficiency in processing large datasets. By implementing reusable question sets within the API, the developer argues that batch processing for mass inquiries—such as evaluating over a million documents monthly—could become significantly more cost-effective. Currently, their use case requires approximately 88% of input tokens for questions, making each interaction with the model expensive, although the total monthly cost remains under $150. Prompt caching would streamline operations by allowing the reuse of question sets without resending them for each request, which could lead to better performance and reduced costs for high-volume applications. The developer highlights potential benefits such as increased accuracy and more complex question sets without substantial retraining efforts. This enhancement could attract wider adoption within the AI/ML community, particularly for projects needing to evaluate large amounts of varied content quickly and efficiently.
Loading comments...
loading comments...