r/snowflake • u/SufficientRelief9615 • 8h ago
Best practice for Cortex Agent token and time budget?
Hello :)
I am currently setting up a Cortex Agent with the aim of using the Cortex Analyst to address complicated text-to-SQL/business questions.
At this point, I would like to establish common limitations of the orchestration budget with regards to the following factors:
– Tokens;
– Seconds.
The Snowflake solution has recommended a time budget of 5 minutes, which was unexpected.
At the moment, I am estimating something of around:
Budget: tokens: 16000, seconds: 300
The objective is primarily to protect against any unusual long reasoning or looping that would lead to unnecessary costs while avoiding rejection of valid questions.
To those who work with Cortex Agents on a day-to-day basis:
What limitations do you normally apply when setting tokens and time?
Do you use fixed limitations or determine precise limitations on the basis of the actual usage, such as P95 + certain margin?
Also, what experience do you have regarding reliability of the mechanism when applying time limits below 5 minutes?
Thanks for the feedback !