OpenAI Prompt Cache Monitoring. A worked example using Python and the… | by Thomas Reid

A labored instance utilizing Python and the chat completion API

As a part of their current DEV Day presentation, OpenAI introduced that Immediate Caching was now obtainable for numerous fashions. On the time of writing, these fashions had been:-

GPT-4o, GPT-4o mini, o1-preview and o1-mini, in addition to fine-tuned variations of these fashions.

This information shouldn’t be underestimated, as it is going to permit builders to avoid wasting on prices and cut back software runtime latency.

API calls to supported fashions will mechanically profit from Immediate Caching on prompts longer than 1,024 tokens. The API caches the longest prefix of a immediate that has been beforehand computed, beginning at 1,024 tokens and rising in 128-token increments. For those who reuse prompts with widespread prefixes, OpenAI will mechanically apply the Immediate Caching low cost with out requiring you to vary your API integration.

As an OpenAI API developer, the one factor you will have to fret about is methods to monitor your Immediate Caching use, i.e. examine that it’s being utilized.

On this article, I’ll present you ways to try this utilizing Python, a Jupyter Pocket book and a chat completion instance.

Set up WSL2 Ubuntu

Source link

The Invisible Revolution: How Vectors Are (Re)defining Business Success | by Felix Schmidt | Jan, 2025

Great Books for AI Engineering. 10 books with valuable insights about… | by Duncan McKinnon | Jan, 2025

AI Ethics for the Everyday User — Why Should You Care? | by Murtaza Ali | Jan, 2025

Despite return, Rams should still prepare for future without Stafford

New Coin Listing – Sealana Crypto Presale Hits $5 Million, 24 Hours Left

Financial Peace University vs. True Financial Freedom vs. Crown Financial MoneyLife

Nigeria not an easy place for startups

Best AI Nude Generators Revealed (2024)

Our Picks

Palestinian journalist Bisan Owda and AJ+ win Emmy for Gaza war documentary | Israel-Palestine conflict News

5 Key Lessons from a Sleep Tech CEO on Brand Building

“Dragonfly Apocalypse” — Thousands of Dragonflies Swarm Beachgoers at Rhode Island’s Misquamicut Beach | The Gateway Pundit

Most Popular

Despite return, Rams should still prepare for future without Stafford

New Coin Listing – Sealana Crypto Presale Hits $5 Million, 24 Hours Left

Financial Peace University vs. True Financial Freedom vs. Crown Financial MoneyLife

OpenAI Prompt Cache Monitoring. A worked example using Python and the… | by Thomas Reid | Dec, 2024

A labored instance utilizing Python and the chat completion API

Set up WSL2 Ubuntu

Related Posts