I have similar concerns about this. The past year or so, software engineers have been encouraged by employers to adopt AI tooling and agentic coding. Now, many enterprises, including mine, are starting to crack down on token spend.
I’ve learned how to fully embrace agentic tooling to do tasks like keeping up with vulnerability reports, initially triaging defects that come in, etc. It would be hard for me to do things the old way at this point when I know tools are available that could make me more productive for a particular category of tasks.
My employer is considered capping all engineers at $200 or $500/mo of token spend depending on level. I regularly spend over $1k/mo today, but believe I can make a strong business justification for the value those tokens are creating.
At this point, I think engineers may be asking what token budgets are when considering new roles.
It's also kind of odd because even a junior engineer has to cost a company around 10 grand a month in salary, taxes, and other costs. So if the company thinks it makes you 10% more effective it seems like it should be an easy choice. So either the companies are shooting themselves in the foot limiting spending or they aren't seeing the productivity boost.
They were talking about what it costs the company, not just the salary. There are taxes, insurances, equipment costs, etc.
10k might be a bit too high, but it's far from "absolutely ridiculous" amounts of being too high. An employee with a 5k salary can easily cost the employer 7-8k
In the US, it’s really difficult to live a middle class lifestyle if you’re not making at least low to mid 100k. I don’t see any starting software salaries at good companies below 100k even in the Midwest.
Sure in the US. But people are throwing token caps in 500, 1000, 2000 dollar range here. That is a very significant expense on top of salary, taxes and benefis. It's only justifiable if they observe 20% 30% or more productivity increases across the board.
If you have huge salaries, yeah sure 500 dollars its okay even if you only get 5 10% more productive.
Fully loaded employee cost is typically around 2x their actual salary. Junior employees making only $60k/yr is absolutely nothing. It’s lowish even in most parts of the western world.
You can also get cheap devs in Central or South America or Eastern Europe too. Doesn’t make it untrue that large parts of the world have junior devs much more expensive than that.
The software development world and ecosystem. I’d say all of the US, major parts of Canada, and major parts of Europe constitute a large part of software developers.
Maybe a dumb question, but what are you doing/seeing others do that uses that many tokens? I find myself pressing the weekly caps on the $20/month plan only when I'm really having the LLMs go wild with the Xtra high effort on the newest/biggest models on abstract problems or where I don't really understand the problem well. (E.g. improving ML models, planning creative molecule synthesis pipelines, identifying subtle performance problems etc)
Generic CRUD/Config stuff I do for work (and assume most software jobs entail? Maybe not correct) doesn't use a significant amount of tokens.
At least for myself personally: Opus 5 Ultra spreading out hundreds of subagents on specific tasks can gobble tokens like mad. Recently I did a custom OCR pipeline in a day and some Revit automation, extremely large datasets, but I definitely spent good money to make it happen so fast.
I almost certainly wasted tokens, but right now things are subsidized, so it works.
I routinely hit the limits on a Claude pro max 20 account for personal projects. For context I have 26 years of professional experience and have been programming for around 40 years.
In addition I have virtually unconstrained usage at work (for now).
Tasks that I do on my personal account:
- building stuff for my wife's business
- analyzing and experimentation and building harnesses to test other models/services
- processing interesting research papers that publish without code or sufficient data to replicate the work
And that's besides simple stuff like exploring new topics I am interested in, and building learning assistants and personal tooling, and eliminating technical chores.
> Maybe a dumb question, but what are you doing/seeing others do that uses that many tokens? I find myself pressing the weekly caps on the $20/month plan
I don’t run into many issue with my personal subscription plans. However, at work while companies like Anthropic offer “Enterprise” plans, you still get billed at API rates which are quite high. The subscription plans are an incredible deal and subsidize these tokens, but these plans aren’t offered to large organizations.
On business side orgs are now absorbing costs of tokens to their own local budgets...so the calculus soon will be:1 extra engineer or extra token spend for existing ones. Tbh, when the productivity payoff becomes more obvious I would happily take the latter.
I tend to think this will end up being a good thing. It will force people to actually think about the best use cases and utilization of tokens while others will push to optimize models and tools to improve cost and efficiency. “Life finds a way”
Yeah but energy is still a cost, and local inference without batch and multiplexing for many users (like an office would be) is even less optimal.
I'm not sure I think it's just pushing the problem forward
You know what happened the last years when the internet went down and there were no emails, no Microsoft Teams, no "npm install", no stackoverflow, no Google, no coordination with other locations?
You know what happened when a PC broke? You know what happens at a construction site when the excavator breaks down?
Rationing tokens on a spend level or count makes no sense to me. Why would I pay $10-20K/mo to employ a software engineer and then balk at a $500, $1000, or even $2000 monthly AI bill, assuming they were even vaguely trying to use the tokens productively?
I’m not an AI-maximalist, but “work a few days with AI and the rest of the month without because of cost” sounds literally crazy to me. (If you think AI is low/zero/negative net value, don’t do the first half; if it has net value, don’t do the second half.)
Still doesn’t make sense if you believe that AI is making a meaningful boost to productivity.
If you are the kind of person who thinks AI will drive down the wages of software engineers that much, you have to think that it is at least a 2x multiplier.
$65k salary all in for a company is easily $100k. So spending anything less than $100k would be what you would aim for.
this not making any sense to you is exactly why companies are implementing these caps: you lack the business acumen to understand how financial resources should actually be deployed.
Please consider rewriting this in a way that does not seem to unnecessarily attack parent. You are not wrong about the incentives, but you may not be right about the execution of the restrictions. I should know. I am living through it.
you're fully correct and I apologize for my tone. similarly to many here, I got hit by the caps and frustrated by how many times coworkers were unintentionally or not burning through these resources.
Claiming you have no idea why the managers had to do this felt to me like avoiding any responsibility
I am the manager in this case and the one who sets the AI policy for my low-8-figure/year product development teams.
Our policy is quite permissive for exactly the reasons I conveyed above. It takes an extremely finely-tuned sense of value and high assurance that you’re on the diminishing return portion of the curve to conclude “SWEs should have this precisely metered amount (rather than zero or ‘as much as they don’t waste’)”
My claiming to have no idea was not an abdication of responsibility but rather a claim that it was likely an error (assuming a non-trivial team size and a business that can continue to grow).
What about a $10k to $20k bill? Or better, 100%-200% of the employees salary? At a certain point in time, finance is going to care greatly about the ability to predict costs. Pre AI, predicting the cost an engineer was quite easy. Salary + benefits + licenses for software. CI systems that were unbounded were rolled up into opex and treated as a separate line item. Where does token spend land in this now?
Finance teams above all else value predictability. Token spend is the opposite of predictable. Eventually these two forces are going to collide and some part of the system will break. My guess is it's token spend, not finance allowing unbound spend on a previously predictable part of the balance sheet.
> Why would I pay $10-20K/mo to employ a software engineer and then balk at a $500, $1000, or even $2000 monthly AI bill, assuming they were even vaguely trying to use the tokens productively?
No idea, but companies refusing to buy a 2000$ laptop once every 3 years, a new mouse, keyboard, or screen for < 500$ are very common, more common than the other type.
So I'd expect those to not suddenly change into "let's waste money" companies.
> Why would I pay $10-20K/mo to employ a software engineer and then balk at a $500, $1000, or even $2000 monthly AI bill, assuming they were even vaguely trying to use the tokens productively?
Spending $10k and then balking at an additional $2k can make sense to me based on the circumstances. It's an additional 20% that hasn't been budgeted for.
The question is partly about what that 20% gives the employer - for some businesses, you might see diminishing returns from additional productivity of a software developer. E.g. I did some work creating BI reports and warehousing for a manufacturing company. Being more productive just meant faster reports or more reports. But faster/more reports didn't necessarily equate to more money for the business.
Sometimes it's a question of a business being penny wise and pound foolish, but sometimes it does make sense to me to not spend this additional money.
If I had to pay $10k salary for a certain job previously, but I now have to pay $10k + $2k for the same job. Why would I not care about that? You were doing your job perfectly well previously, why do you want more budget?
Because there are diminishing returns somewhere for sure. Will the developer be more effective when he gets $2,000 more in tokens? I can't max the $100 subscription. The same for any other employee-related expense. It's a no-brainer to give the developer a solid laptop. But if you'd pay extra $2,000/month in computer and office supplies costs, will it save the employee a few hours? I don't think so. Their salary doesn't matter that much.
This sounds roughly similar to, "If eating steak is nutritious and tastes good, why would you not eat an 8oz sirloin for breakfast, lunch, and dinner for an entire month?"
"Too much of a good thing" does exist. Someone might feel that it takes time for their organization to "metabolize" the rapid additions made by coding agents. That seems quite reasonable to me. Your comment seems to argue a false dichotomy.
Forward thinking employers should use these limits with a holistic perspective. It's obviously time to take a deserved coffee break if you've token maxxed too much. Maybe take Friday off, AI now allows us to accomplish in four working days what used to take five.
I don't think people understand, tokens must never be constrained.
Right now, LLMs are the worst they'll ever be, but imagine what they will be like at the peak: Anything you want to code, coded instantly. Not waiting for code to stream in or wait hours for some agents to crunch through loops and planned: it just appears on the screen instantly like the way a webpage loads. Then imagine it can be done locally, on your device. Need an entire new custom operating system from scratch for some obscure hardware? Done. Here it is.
That's going to be like pure crack to anyone who needs to do absolutely anything. That's going to be like our generation's version of "today's supercomputers will someday be in everyone's pocket".
Right, every technology is suboptimal in the beginning, but constraints will ever exists, energy and hardware in primis. Aldo you have to factor entropy in regardless, complexity remains a cost
Imagine if people had said one day computer programs had to be more memory constrained because everything uses too much RAM. Madness. We'd never see 128GB of RAM.
This is actually how Jeff Dean, Chief Scientist at Google recommends doing thought exercises to challenge assumptions (watch is recent YX Startup School talk). The end goal isn’t “128GB of RAM”, the goal is “useful effective software”. If for whatever reason RAM was an immovable constraint, a different tech tree would emerge.
The tech tree where we develop tech to have tons of RAM is better and leads to more opportunities than the tech tree where everything is built for only a handful of GB of RAM. We may never have had LLMs if we didn’t go down the first path.
Or maybe we’d have LLMs that fit in a handful of GBs (and we technically already do). It’s hard to tell without completely going down that parallel universe and simulating all the milestones in it, but it’s a fun thought experiment.
We would be lot better off. I am very weirded that computers are one of the place where waste is sometimes even celebrated. We would be laughed out if we suggested same with say energy in general. Why have walls, we could just have enough space heaters to always push warm air or cool air... Saying we need walls stops us reaching point where we heat or cool ourselves in outdoors constantly.
if i’m not mistaken, in the past, people had limited time on “the mainframe”. so eventually they brought in what would be the equivalence of a local model… small computers that could do smaller tasks locally so they didn’t have to keep shelling out money to the mainframe gods and be handcuffed for usable time.
i could absolutely be mistaken that this is how it worked, it was before my time. but this is how people explain it worked for them.
i don’t think most work gives a shit about soa. smaller repeatable tasks can absolutely be run just fine on smaller local models. sure, we’ll upgrade models occasionally just like we went from suitcase sized laptops to whatever we use today.
this idea the hypedorks are pushing that soa is the only way is hilarious. hobbyists spend stupid money on classic cars, tools for woodworking, or whatever. spending money to do a hobby at home has never stopped wonks and their hobbies and businesses will do the same, spend to run models locally.
the sooner the hypeTrash does what hype always does, fades to irrelevance, the sooner we can get back to work.
In my personal anecdotes, Luna has changed calculations, once again; I'm _almost_ unconstrained in spending (my employer's Copilot subscription) tokens as long as its Luna. And it works more than good enough.
I’m using Luna all the time now since it seems quite good enough for everyday programming and so that my $20/month ChatGPT subscription doesn’t hit the weekly limit.
It’s an artificial constraint. I could afford to spend more on my hobbyist programming but I choose not to.
I suspect that I’m adapting to the model by giving it more concrete guidance and guardrails. I read code more and I tell it to refactor code that I don’t like. Working within constraints isn’t all bad.
On-prem AI is another solution. I’m surprised more companies aren’t leaning in this direction due to security/IP concerns. My employer is planning to spend several million on local AI hardware.
I think we're still in the early days of figuring out how to use AI as a daily work tool. It's going to take time to develop the right practices. I do see the point in having token limits, they force you to be intentional.
My guess is that eventually we'll end up with small, local LLMs that are unlimited for routine stuff, and we'll save the big expensive models (with quotas) for the heavy lifting.
Currently navigating this situation at my employer. We went from virtually unlimited token spend per software engineer, to $150 per month due the recent change in billing terms from GitHub Copilot. Rationed over a month, about $7.50 per day. Basically, a couple bad apples spoiled the bunch (contractors using Opus to center divs).
We're 15 days into this new policy and its going ~okay~. Engineers adapt as they do, and have been leaning on `gpt-5.6-luna xhigh`. Some contractors have already hit their budget limit for the month.
A couple observations here:
1. Because LLMs/agents are tools, limiting their usage becomes a distraction and ends up being more of a drag on each individual's productivity. Instead of "just doing the work" engineers are now wasting time tinkering with setups (graphs, caveman skills, etc.).
2. Skill atrophy is real. Engineers that hit their budgets are seemingly less productive and less capable which is deeply concerning.
3. There are legitimate conversations happening now to explore open source harnesses (opencode/pi) and open weight models at the company in order offset the costs associated with going through a standard provider.
4. Token prices are venture capital subsidies. Its essentially free samples to get the market hooked on their addictive white collar drug. Remember when an Uber cost $7? As soon as OpenAI and Anthropic go public, they will need to begin showing progress toward profitability. That is when the true price of a token will be revealed.
I don't even understand how it's possible to max out these subscriptions. In the last few weeks I've been basically vibecoding at work with 5.6 Sol at max thinking and max context, ~6hrs per day of the thing chugging along, sometimes 2-3 at a time. And it's not even a quarter used up.
I know some people are doing 10-agent parallel harness stuff, but that cannot possibly be the norm, and I find it hard to believe any company is expecting that level of output.
cebert | 6 hours ago
I’ve learned how to fully embrace agentic tooling to do tasks like keeping up with vulnerability reports, initially triaging defects that come in, etc. It would be hard for me to do things the old way at this point when I know tools are available that could make me more productive for a particular category of tasks.
My employer is considered capping all engineers at $200 or $500/mo of token spend depending on level. I regularly spend over $1k/mo today, but believe I can make a strong business justification for the value those tokens are creating.
At this point, I think engineers may be asking what token budgets are when considering new roles.
alainrk | 6 hours ago
theChris-in | 6 hours ago
neaden | 6 hours ago
andriy_koval | 6 hours ago
pdhborges | 6 hours ago
neaden | 6 hours ago
izacus | 6 hours ago
jdrek1 | 6 hours ago
10k might be a bit too high, but it's far from "absolutely ridiculous" amounts of being too high. An employee with a 5k salary can easily cost the employer 7-8k
pdhborges | 5 hours ago
cebert | 5 hours ago
pdhborges | 5 hours ago
If you have huge salaries, yeah sure 500 dollars its okay even if you only get 5 10% more productive.
Kirby64 | 6 hours ago
pdhborges | 5 hours ago
cebert | 5 hours ago
Kirby64 | 4 hours ago
pdhborges | 3 hours ago
Kirby64 | 3 hours ago
The software development world and ecosystem. I’d say all of the US, major parts of Canada, and major parts of Europe constitute a large part of software developers.
the__alchemist | 6 hours ago
Generic CRUD/Config stuff I do for work (and assume most software jobs entail? Maybe not correct) doesn't use a significant amount of tokens.
unified101 | 6 hours ago
the__alchemist | 6 hours ago
goolz | 5 hours ago
I almost certainly wasted tokens, but right now things are subsidized, so it works.
ygjb | 5 hours ago
In addition I have virtually unconstrained usage at work (for now).
Tasks that I do on my personal account: - building stuff for my wife's business - analyzing and experimentation and building harnesses to test other models/services - processing interesting research papers that publish without code or sufficient data to replicate the work
And that's besides simple stuff like exploring new topics I am interested in, and building learning assistants and personal tooling, and eliminating technical chores.
cebert | 5 hours ago
I don’t run into many issue with my personal subscription plans. However, at work while companies like Anthropic offer “Enterprise” plans, you still get billed at API rates which are quite high. The subscription plans are an incredible deal and subsidize these tokens, but these plans aren’t offered to large organizations.
monkeydust | 6 hours ago
Aboutplants | 6 hours ago
deadbabe | 6 hours ago
alainrk | 6 hours ago
deadbabe | 6 hours ago
yladiz | 5 hours ago
deadbabe | 3 hours ago
Not even for the environmental aspects, just purely from energy demands, at any spot on the Earth where a data center could go.
stephbook | 6 hours ago
You know what happened when a PC broke? You know what happens at a construction site when the excavator breaks down?
Exactly nothing — and that's okay.
sokoloff | 6 hours ago
I’m not an AI-maximalist, but “work a few days with AI and the rest of the month without because of cost” sounds literally crazy to me. (If you think AI is low/zero/negative net value, don’t do the first half; if it has net value, don’t do the second half.)
Avicebron | 6 hours ago
iugtmkbdfil834 | 6 hours ago
sarchertech | 5 hours ago
sarchertech | 5 hours ago
If you are the kind of person who thinks AI will drive down the wages of software engineers that much, you have to think that it is at least a 2x multiplier.
$65k salary all in for a company is easily $100k. So spending anything less than $100k would be what you would aim for.
loki-ai | 6 hours ago
iugtmkbdfil834 | 5 hours ago
loki-ai | 5 hours ago
Claiming you have no idea why the managers had to do this felt to me like avoiding any responsibility
sokoloff | 5 hours ago
Our policy is quite permissive for exactly the reasons I conveyed above. It takes an extremely finely-tuned sense of value and high assurance that you’re on the diminishing return portion of the curve to conclude “SWEs should have this precisely metered amount (rather than zero or ‘as much as they don’t waste’)”
My claiming to have no idea was not an abdication of responsibility but rather a claim that it was likely an error (assuming a non-trivial team size and a business that can continue to grow).
abuani | 6 hours ago
Finance teams above all else value predictability. Token spend is the opposite of predictable. Eventually these two forces are going to collide and some part of the system will break. My guess is it's token spend, not finance allowing unbound spend on a previously predictable part of the balance sheet.
gibolt | 5 hours ago
delusional | 5 hours ago
What if they feel like they are getting value from it, but are actually not?
izacus | 6 hours ago
No idea, but companies refusing to buy a 2000$ laptop once every 3 years, a new mouse, keyboard, or screen for < 500$ are very common, more common than the other type.
So I'd expect those to not suddenly change into "let's waste money" companies.
verbify | 5 hours ago
Spending $10k and then balking at an additional $2k can make sense to me based on the circumstances. It's an additional 20% that hasn't been budgeted for.
The question is partly about what that 20% gives the employer - for some businesses, you might see diminishing returns from additional productivity of a software developer. E.g. I did some work creating BI reports and warehousing for a manufacturing company. Being more productive just meant faster reports or more reports. But faster/more reports didn't necessarily equate to more money for the business.
Sometimes it's a question of a business being penny wise and pound foolish, but sometimes it does make sense to me to not spend this additional money.
delusional | 5 hours ago
krab | 5 hours ago
cfiggers | 5 hours ago
"Too much of a good thing" does exist. Someone might feel that it takes time for their organization to "metabolize" the rapid additions made by coding agents. That seems quite reasonable to me. Your comment seems to argue a false dichotomy.
veeti | 4 hours ago
deadbabe | 6 hours ago
Right now, LLMs are the worst they'll ever be, but imagine what they will be like at the peak: Anything you want to code, coded instantly. Not waiting for code to stream in or wait hours for some agents to crunch through loops and planned: it just appears on the screen instantly like the way a webpage loads. Then imagine it can be done locally, on your device. Need an entire new custom operating system from scratch for some obscure hardware? Done. Here it is.
That's going to be like pure crack to anyone who needs to do absolutely anything. That's going to be like our generation's version of "today's supercomputers will someday be in everyone's pocket".
alainrk | 6 hours ago
deadbabe | 6 hours ago
jtfrench | 5 hours ago
deadbabe | 3 hours ago
jtfrench | 2 hours ago
deadbabe | an hour ago
Ekaros | 5 hours ago
toofy | 6 hours ago
if i’m not mistaken, in the past, people had limited time on “the mainframe”. so eventually they brought in what would be the equivalence of a local model… small computers that could do smaller tasks locally so they didn’t have to keep shelling out money to the mainframe gods and be handcuffed for usable time.
i could absolutely be mistaken that this is how it worked, it was before my time. but this is how people explain it worked for them.
i don’t think most work gives a shit about soa. smaller repeatable tasks can absolutely be run just fine on smaller local models. sure, we’ll upgrade models occasionally just like we went from suitcase sized laptops to whatever we use today.
this idea the hypedorks are pushing that soa is the only way is hilarious. hobbyists spend stupid money on classic cars, tools for woodworking, or whatever. spending money to do a hobby at home has never stopped wonks and their hobbies and businesses will do the same, spend to run models locally.
the sooner the hypeTrash does what hype always does, fades to irrelevance, the sooner we can get back to work.
theanonymousone | 6 hours ago
unified101 | 6 hours ago
Tomorrow 2 engineers at 50% token time will change to 1 engineer with 100% token time for the same work.. the latter is too cost effective.
skybrian | 5 hours ago
It’s an artificial constraint. I could afford to spend more on my hobbyist programming but I choose not to.
I suspect that I’m adapting to the model by giving it more concrete guidance and guardrails. I read code more and I tell it to refactor code that I don’t like. Working within constraints isn’t all bad.
Taikhoom2010 | 5 hours ago
https://s-1.vercel.app/posts/the-struggle-of-openai/
variadix | 5 hours ago
ironqcold | 5 hours ago
jgmedr | 5 hours ago
We're 15 days into this new policy and its going ~okay~. Engineers adapt as they do, and have been leaning on `gpt-5.6-luna xhigh`. Some contractors have already hit their budget limit for the month.
A couple observations here:
1. Because LLMs/agents are tools, limiting their usage becomes a distraction and ends up being more of a drag on each individual's productivity. Instead of "just doing the work" engineers are now wasting time tinkering with setups (graphs, caveman skills, etc.).
2. Skill atrophy is real. Engineers that hit their budgets are seemingly less productive and less capable which is deeply concerning.
3. There are legitimate conversations happening now to explore open source harnesses (opencode/pi) and open weight models at the company in order offset the costs associated with going through a standard provider.
4. Token prices are venture capital subsidies. Its essentially free samples to get the market hooked on their addictive white collar drug. Remember when an Uber cost $7? As soon as OpenAI and Anthropic go public, they will need to begin showing progress toward profitability. That is when the true price of a token will be revealed.
eudamoniac | 2 hours ago
I know some people are doing 10-agent parallel harness stuff, but that cannot possibly be the norm, and I find it hard to believe any company is expecting that level of output.
[OP] fullautomation | 5 hours ago