I agree with your observation. Although the framing of the question signals some of the response "Do you think.. " forces the model to check for the cost-benefits.
We can definitely mould AI agents to think more critically about these things, I dont know how effective it will be in the long term honestly.
Recommend checking out Ponytail — a pi extension that behaves like a senior engineer who keeps the LLM code generation in check, and the processing of your prompts just the same
We can definitely mould AI agents to think more critically about these things, I dont know how effective it will be in the long term honestly.