it really depends on the model. best rn is opus 4.6 (i will never pay out of my own pocket to scummy ai companies so i understand not wanting to try it)
they get 95% of the way there, and the last tuning part is the hardest
using subagents and letting it iterate itself (especially if u have testcases that it can use to verify itself) is handy
for me, its useful because work and my hobby (which is linux mobile) are diff. i can parallelize my work and enjoy other life too