Rendered at 18:45:28 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
MiddleEndian 22 hours ago [-]
This article is a bit rambly so I'll just focus on some things from the beginning:
>If your kitchen knife kept changing shape, weight, and edge, you’d have to relearn it every time; that’s a hard tool to build trust in.
This concept was betrayed far before agentic tools, with a much earlier concept: Automatic updates.
To use one product as an example: When Windows ME and Windows Vista came out, people hated them even more than they usually hated Windows, so they did not use them. Microsoft was forced to respond by making a not-quite-as-bad OS in Windows XP and a pretty good OS in Windows 7 respectively. No longer is that an option, your workflow will simply be interrupted by automatic updates.
>Vim and Emacs, in their infinite customizability, can be molded to fit your exact hand and workflow
Vim is one major exception to the automatic update problem. I trust vim not just because it can do a ton of shit (although that is certainly nice), but because unlike most other software, its UI doesn't change unless I tell it to change. Aside from switching from vim to neovim (my decision, not a forced update), my muscle memory from a couple decades ago still works today.
ozim 20 hours ago [-]
That's not just developers as you noted.
It is "velocity fallacy" — product people want "all the features ASAP or right away".
Until users with their managers come with pitchforks and torches. I worked on such internal project where we as developers were able to deliver new features and new version every 2 weeks (which is not a pinnacle of the game of course) and were thinking if we can move to daily delivery. Because we were good devs and wanted to appease product owner.
Until one day product owner came back with feedback, how everyone is pissed off at him for shifting ground under people feet, while he also thought he is doing great delivering all those great features. It was pushed back to limited amount of features each month.
People need training, people need to understand what and why.
In the end it turns out it is also AI coding fallacy, because most of the software is built for limited audience, which has its specific timeline on accepting and internalising new features or different ways of doing stuff. Unless we take humans from the equation and we start building for AI itself.
BobbyTables2 19 hours ago [-]
I’ve seen this too and it’s darkly humorous.
We published 1-2 releases of our component each month. Eventually another (internal) team would pick it up along with others, test, and release the combined set maybe 1-2 times a year.
Customers being extremely risk adverse never wanted to update, due to risk of changes combined with the interruption, even though we were fixing serious bugs left and right from the earlier rushed development.
We’d get escalations on things that were fixed years ago.
Probably should have spent another 3-5 years getting the first release in better shape instead of spending 3-5 years flinging flaming turds to the paying customers.
The executives slowly released they inexplicably lost half the business compared to the previous generation as upstart competitors stole their market share.
You know that famous “how shit happens” tale? I’m pretty sure each layer of management was telling the next how powerful the product was… few could abide by it.
pmontra 9 hours ago [-]
I'm not sure that I fully understood but:
> Probably should have spent another 3-5 years getting the first release in better shape instead of spending 3-5 years flinging flaming turds to the paying customers.
I wonder if the company would have had the money to pay salaries for 3 years, unless for those "paying customers" that obviously started paying much earlier than that.
And
> The executives slowly released they inexplicably lost half the business compared to the previous generation as upstart competitors stole their market share.
s/released/realized/ ? Exactly one of my worries. Maybe they could have created those competitors themselves. Two brands, one for the original product and one for a product with a somewhat different layout and all the improvements that they did not dare to add the the original product. But then you need a third brand, a fourth one, etc.
ent001 6 hours ago [-]
I think software is just very young. We've been doing it for maybe 2 generations. The current way, saas, ci/cd, is like half a generation old.
We're still at the "search" phase, there is very little wisdom at social and personal levels.
It's clear from the rapid flux in tooling, methods, ideology, and the results, that we basically have no idea what we're doing re: using computers and building software. A lot of the self assuredness of current advice is self soothing behavior.
kelnos 17 hours ago [-]
> Customers...never wanted to update, due to risk of changes [...] We’d get escalations on things that were fixed years ago [...] as upstart competitors stole [our] market share.
This is so weird to me, though. Your customers had two choices:
1. Stay with your software and upgrade to the latest version, where the bugs they were hitting were fixed, and risk some amount of retraining due to UI/UX changes.
2. Switch to a completely new, different product, which guarantees retraining (possibly a lot more), and probably would contain the same or different bugs they were fighting with when using your software.
And... they went with #2?? I get that people think the grass is always greener on the other side, but there's a reason why we have that saying!
BobbyTables2 17 hours ago [-]
I suspect they were thrilled at the prospect of a better made product (not encumbered by historical baggage) that cost less, had a lower friction sales process, made by a company actively trying to win them over.
pjmlp 12 hours ago [-]
Yes, they go with #2 because they are really p***.
As anecdote of one, I was once part of a project to port a .NET Framework to Java, because the customer was really annoyed with the rewrite, as the application relied heavily on the .NET Features that never made the cut to modern .NET.
Another two .NET heavy weights in .NET CMS space, Sitecore and Optimizely, nowadays rely on JS/TS frameworks for their extensibility SDKs on the SaaS products for headless deployment, only the classical (older) PaaS still support .NET as extension language.
3eb7988a1663 16 hours ago [-]
You are assuming the changes would all be positive, minus some UI hassle. I have seen plenty of products where vital workflows were scrapped entirely or made significantly more clunky.
Modern software is quicksand. Maybe an update improves things, but users have all experienced working software made worse.
TeMPOraL 12 hours ago [-]
Yup. Powerful features are removed to "streamline experience". Sometimes in bulk, because product migrated from e.g. bespoke native app to multiplatform webshit in an embedded browser, which automatically halved the performance and removed most ergonomics.
And then new features are released as MVPs - meaning, they do the absolute minimum to check the box (almost a literal box: close the feature ticket internally), and inefficiently so. Whether it'll be iterated on afterwards, depends. Pretty unlikely in the immediate term. They need to "collect usage data to know what to do" first, which means it gets deprioritized relative to moar features.
But that's fine for me. As long as copy-paste works, I can make do with other software, perhaps competing software, or worst case, have Claude find and use some powerful-but-janky OSS CLI tool to do the stuff for me.
I'm less angry at it than I used to. Mostly because I don't have time to be annoyed anymore, but it's true that I've learned to like updates from few companies. That's because after years - years - I've noticed things gradually improving on average. True of Android & Samsung OneUI, except when it's not. True of UniFi stuff. If I see some new feature behaving badly, I now mostly trust they'll eventually fix it. It'll take a year or three, they'll overhaul it entirely twice, and I may need to buy a new phone to get it, but it will happen one day. But most software doesn't even clear that bar.
pjc50 9 hours ago [-]
"Trust thermocline". This happens when people no longer trust that improvements are going to come from #1.
Someone whose partner is about to leave them loudly proclaims "I've changed!". Should they believe them? Often the answer turns out to be "no".
It's quite difficult to actually build back trust in these situations.
mendapi 17 hours ago [-]
[flagged]
hilariously 20 hours ago [-]
This is exactly what I said to the last company (which is basically shutting down now lol) management - "infinite customizeability" means "zero percent understandability" for a typical suite of software.
I literally was told not to narrow down the features too much or the product would be too constrained... well then wtf does it actually _do_ for anyone?
kelnos 17 hours ago [-]
I don't think infinite customizeability is a problem. Poor default configurations are the problem when you build in the ability to customize.
The idea is that the user should be able to be productive more or less immediately after installing the software, and then incrementally customize it as they discover they have different needs than the defaults give them.
A lot of developers/companies forget the first part, and ship something that requires a team of consultants six months to set up before anyone can do anything useful with it. Of course no one wants that. (Well, except for the highly-paid consultants.)
3eb7988a1663 16 hours ago [-]
What mythical software is this? The only infinitely customizable software I can think of is emacs, and that takes an enormous amount of effort to become proficient.
ssdspoimdsjvv 7 hours ago [-]
any serious enterprise software that attempts to fit every company's workflow (rather than forcing the company to follow theirs)
Scarblac 13 hours ago [-]
Spreadsheet software, like Excel.
Animats 14 hours ago [-]
> I don't think infinite customizeability is a problem. Poor default configurations are the problem when you build in the ability to customize.
This is a classic IDE problem. The IDE has a preferred directory layout it wants, and some other build tool you're using has a different directory layout.
_carbyau_ 12 hours ago [-]
Sensible defaults and/or a basic wizard to choose templated configs or assist with customising the few things that have to be.
Scarblac 13 hours ago [-]
OTOH, there's spreadsheets. The abstraction is just so great that users can and do use them for anything.
The problem seems to be that there's not many such great abstractions to be found.
ozim 12 hours ago [-]
But parent is still correct on "zero percent understandability".
I have seen dozens upon dozens of Excel sheets which "just worked" until they didn't and then of course person who could fix that left company 10 years ago.
Besides I also know dozens of people whose life would be easier if they would learn a bit more of Excel like pivot table is there for them.
hilariously 8 hours ago [-]
For some quantity of use. When I say infinite customizeability I mean its just fronting chat gpt + tools without any real product use case.
saimiam 18 hours ago [-]
Just up thread from your comment, someone commented in a positive way on the infinite customisability of vi.
So which is preferable - customisability or very limited feature set?
ozim 11 hours ago [-]
Question is: what are you building?
There is no general answer that will answer which is preferable.
It is easy for people to come up with features they want or a customisation and they don't care for the cost of actually maintaining that feature or customisation.
I work on building SaaS platform, we had multiple customers for whom we build custom features and they paid for development of those features. Not fun part is after 2-3 years those customers are gone (for example employees at customer rotated and they switched to something totally different) — but now we are left with dead feature no one is paying for supporting, some are easy to remove, some are not.
kelnos 17 hours ago [-]
I think the problem with the vi/vim example is that it's very much not for everyone. I use it, and have used it for 25-odd years, but I'm not an expert at it by any means (because I haven't taken the time to try to be), and occasionally have to look up how to do something. But I've customized it to my heart's content, and feel very productive in it.
But I would not expect a large percentage of developers to want to do that. I've used IDEs as well, and they're mostly fine, and work really well for lots of people. Many IDEs are pretty damn customizeable too, though most people may not change many settings.
songhonglei1985 16 hours ago [-]
[flagged]
cosmic_cheese 17 hours ago [-]
A big driver of automatic updates in the late 00s and early 10s were people who never updated their OS or browser and were getting compromised left and right. As I recall, one of the earliest widespread users automatic updates was Google with Chrome, which they did because otherwise there was no way to respond to severe 0days and such before huge swathes of users got hit.
So I think there's a place for automatic updates, but the feature should be restricted to security fixes only. Using it to foist ill-conceived changes on users is just abuse.
brabel 10 hours ago [-]
Yeah people never wanted to update Windows since they had been burned in previous upgrades and their current setup just worked. Not only were upgrades difficult to do and took a long time, many times the new version was just worse or requires relearning lots of things. For what, asked the normal user?! But MSFT tied IE browser updates to OS updates just to make things so much worse for those people since now their choice was to either upgrade and have to deal with the distress of your UI changing, or be vulnerable to widely known exploits in your outdated browser.
My late dad preferred to not upgrade at any cost to the end!
cosmic_cheese 6 hours ago [-]
That holds for major upgrades (e.g XP → Vista), but those didn’t used to just happen on their own. In the time frame my previous post was speaking of, one still had to buy major upgrades, and updates within the same version rarely brought significant UI changes. So in reality, there wasn’t much valid reason for there to be for example XP and 7 machines still running initial releases or early service packs for years on end (which was shockingly common). Most of the reason people had for avoiding minor updates is that they just found them annoying.
Now today of course things are quite different and it’s not unusual for a routine Windows update to turn things upside down. Users are more justified in update-averseness than they were 15-20 years ago, except of course now that brings much greater risk of getting pwned than it did back then.
RossBencina 14 hours ago [-]
Agree. Security fixes are the killer app for automatic updates. I would argue there are other kinds of fixes that should make the cut: crash bugs, data loss bugs. Maybe even broken features, interaction annoyances, etc. Then it becomes a slippery slope, and the risk/benefit trade-offs are harder to arbitrate.
I don't think "ill-conceived" is giving enough credit. It has been clear to me for a long time that "security" is the justification for pushing the user to update, but the updates themselves are frequently leveraged as a vector for other, less user-friendly, practices.
BrenBarn 15 hours ago [-]
> A big driver of automatic updates in the late 00s and early 10s were people who never updated their OS or browser and were getting compromised left and right.
I think we would be better off if people worked harder to prevent vulnerabilities before releasing the software.
wolrah 2 hours ago [-]
> I think we would be better off if people worked harder to prevent vulnerabilities before releasing the software.
Pawn Stars meme: "Best I can do is people using glorified chatbots to generate mediocre code an order of magnitude faster"
That said, the vast majority of the exploits that led to the widespread adoption of automatic update mechanisms were based around memory safety bugs which we do in fact have solutions to entirely prevent in most newly developed software these days.
cosmic_cheese 15 hours ago [-]
I won't argue with that, but with how complex web browsers have become, holes are unavoidable regardless of the level of effort put forth to prevent them.
jcranmer 21 hours ago [-]
> Microsoft was forced to respond by making a not-quite-as-bad OS in Windows XP
Prior to XP, MS had two lines of Windows: the Windows 9x kernels and the Windows NT kernels. Windows XP was meant to be the merger of the two lines, adapting Windows NT to have compatibility with Windows 95 and Windows 98 features. Unfortunately, Windows XP development went overlong, so MS wedged in Windows ME to give a stop-gap release until XP could actually be released.
musicale 21 hours ago [-]
> When Windows ME and Windows Vista came out, people hated them even more than they usually hated Windows, so they did not use them. Microsoft was forced to respond by making a not-quite-as-bad OS in Windows XP
This timeline doesn't seem to make sense, as Windows XP came out in 2001, and Windows Vista in 2006-2007. Maybe you are referring to Service Pack 3 in 2008?
ToValueFunfetti 20 hours ago [-]
Reactions to Windows ME forced MS to respond with XP; reactions to Vista forced them to respond with 7. They're being stated in parallel rather than chronologically.
accrual 19 hours ago [-]
This isn't quite right, Windows Me was a stopgap between Windows 98 and Windows XP. Microsoft had no intention of keeping anyone on Me for long. There were already plans to get consumers onto an NT-based OS when Me launched, the initial version of this was codenamed Neptune.
> Microsoft discussed a plan to delay Neptune in favor of an interim OS known as "Asteroid", which would have been an update to Windows 2000 (Windows NT 5.0), and have a consumer-oriented version. At the WinHEC conference on April 7, 1999, Steve Ballmer announced an updated version of Windows 98 known as Windows Millennium, breaking a promise made by Microsoft CEO Bill Gates in 1998 that Windows 98 would be the final consumer-oriented version of Windows to use the MS-DOS architecture. [0]
So XP was not a reaction to Me's reception, it was already in the works as a replacement when Me came out.
I feel like everyone is forgetting Windows 2000, which IMO is the best Windows that MS ever released. The main problem with it was hardware support, since people were hesitant to move away from the 95/98/ME line, and hardware manufacturers were still playing catch-up, as most didn't support NT 3/4 at the time and didn't see a need, as it was largely a server/business OS. But otherwise it was rock solid, fast, and a breath of fresh air.
I did try XP here and there, but was instantly turned off by the cartoonish default theme (yes, I know you could change it). It was buggier than Win2k, and I didn't have the patience to wait around until they fixed it. I was told later on by people in the know that Service Pack 3 was the bees knees, but by then I'd moved on to Linux full-time (with some jaunts into OS X) and felt no need to come back.
pjmlp 12 hours ago [-]
Nope, XP was the introduction of Windows NT linage into mainstream computing, ME was basically 98 with a few goodies to keep selling newer 9x versions in the meantime.
bilkow 20 hours ago [-]
Your quote is missing the end, required for it to make sense (irrelevant parts omitted):
> When Windows ME and Windows Vista came out [...]. Microsoft was forced to respond by making [...] Windows XP and [...] Windows 7 respectively.
It's basically from ME and Vista to XP and 7, respectively. AFAIK respectively in this context means that for ME, they were forced to respond with XP, and for Vista, they were forced to respond with 7.
MiddleEndian 20 hours ago [-]
Yes, this is what I meant.
Also to respond generally to other posts. I am aware of the the separation between NT and 9x. That was Microsoft's problem and not the customers' problem. They could not force people onto ME and had to actually appeal to customers with XP. Then later on, they could not force people onto Vista and had to actually appeal to customers with 7.
Automatic updates remove the step where they have to appeal to anybody.
And I don't mean to single out MS. I remember having a Mac and switching from some version of OS9 back to 8.6 for some reason (don't recall why, but it doesn't matter because it was my computer so it was my decision). Nowadays people are complaining about Liquid Glass and they cannot rollback their OS on their iDevices.
eterm 20 hours ago [-]
Indeed it makes no sense at all, because Windows 2000 even came out before Windows ME.
xnx 5 hours ago [-]
Windows Vista was disliked at launch primarily it ran slowly on existing hardware, suffered from widespread driver incompatibilities, and introduced constant security pop-ups.
UI changes were pretty minor, especially compared to the flat design of windows 8.
GuB-42 19 hours ago [-]
> Vim is one major exception to the automatic update problem.
Except for one thing that pissed me off a great deal. I am not a true vim user, though I use it from time to time, because it is everywhere and it works through ssh. At some point they introduced "visual mode", and it turns on when you start using the mouse, it broke the way I used the mouse copy-paste in vim. I know I can do "set mouse-=a", but when I am just using vim as "the standard text editor" (sorry ed), I don't want to configure anything as it is usually a one shot job.
I never asked for that, at some time, it just happened. I guess as a major upgrade, but the thing is, something changed that I didn't want to change.
I understand the reason for this change, also https://xkcd.com/1172/ but I just wanted to say that even vim is not immune.
Some tools are immune though. Usually command line tools used in scripts. For example "apt-get" doesn't change, but "apt" does. "apt-get" is what you use when you want a stable interface (especially scripts), "apt" is for when you want something nicer.
Izkata 16 hours ago [-]
I think visual mode has existed as long as vim existed, and vi was what didn't have it.
When I first started at my company, we did all work on shared VMs, and the system vimrc had that "mouse" setting enabled. Something an employee had done decades ago to be helpful, really confused me until I realized what was going on. I'm thinking your distro, not vim, is what flipped the setting.
This copies your ~/.vimrc if unmodified to every server you ssh to.
(For bonus points, you can make a program to copy ~/.ssh/config around as well so that your config gets copied to servers you ssh to from there.)
tremon 6 hours ago [-]
> unlike most other software, its UI doesn't change unless I tell it to change
My entire reason for switching from vim to neovim was that vim did change its UI, by neutering /etc/vimrc in a major update a while ago.
denkmoon 19 hours ago [-]
What's really freakin cool about AI development is other people can check rules into the repo which change the shape of my tools :)
kelnos 17 hours ago [-]
That sounds incredibly annoying to me. Not cool at all. (I would assume you're being sarcastic except for the smiley at the end of your comment.)
The thing about it, though, is that LLMs aren't "your" tools. They're someone else's tools, and you are at the whims of day-to-day changes to them.
pmontra 9 hours ago [-]
As a data point: I'm using Claude Pro inside Emacs right now. I can copy/paste better than when I use it inside the terminal. I never installed a custom IDE since maybe 2012 (my last Java project) so I can't really compare the experiences but I don't have anything to complain to the agentic pairing of Claude and Emacs.
thayne 17 hours ago [-]
> its UI doesn't change unless I tell it to change
Vim and emacs aren't unique in this. In fact quite a few non-commercial FOSS projects don't make significant changes to the UI. I think there are (at least) a couple of reasons for that. First, there isn't usually pressure to constantly attract new users, so there isn't as much pressure to change the UI to make it easier or simpler for new users, or follow UI trends. Secondly, the projects often don't have dedicated UI/UX designers who want to try out new ideas or make their mark, etc.
However, these aren't strictly good things, you can end up with an unintuitive UI with a steep learning curve that is hard to learn.
devin 16 hours ago [-]
Your view of what constitutes a tool is overly broad. An OS is not a tool. A hand plane is a tool. It’s sharp and does one thing. The stuff an OS does might enable tools, and it may contain many tools, but it is not itself a tool.
brianpan 16 hours ago [-]
A food processor is a kitchen tool even though it does many things.
A computer is also a tool that does many things and the OS is arguably the important part of how a user wields that tool.
lelanthran 13 hours ago [-]
> A food processor is a kitchen tool even though it does many things.
It does variations of the same thing.
A computer is more of a toolbox than a tool.
inigyou 8 hours ago [-]
Isn't a food processor just what certain marketing teams call a blender?
tremon 6 hours ago [-]
Blending is just one of the functions of a food processor. A food processor can usually also do kneading, mincing, and slicing, or (depending on extensions) even rolling (pasta) dough and squeezing citrus fruits.
connicpu 20 hours ago [-]
This is part of why my text editor is built from source on my own branch where I only occasionally pull in changes from master. I know I'm probably a very rare exception here.
jgord 21 hours ago [-]
A good tool does a job well with minimal side effects and maximal predictability.
Perhaps this is why users dislike monthly SaaS - they cannot trust stability of the tool, because often the incentives are to keep adding features well past peak utility [ resulting in enshitification ]
jgord 21 hours ago [-]
followup .. one reason I now love and detest C++ is the regularity of new features in the core working set [ particularly the current politically correct incarnation of the smart pointer. ]
Razengan 14 hours ago [-]
> If your kitchen knife kept changing shape, weight, and edge, you’d have to relearn it every time; that’s a hard tool to build trust in.
On the other hand, if you hire a cook, then all you need to know is English (or whatever language they speak).
Even if the cooks keep changing, you can always just tell them "use the knife"
peterbower 20 hours ago [-]
Vista came out after XP. They rushed 7 out after that.
devmor 15 hours ago [-]
> This concept was betrayed far before agentic tools, with a much earlier concept: Automatic updates.
This is a huge complaint for me. I had to disable JetBrains from automatic updates because they wont stop trying to force their horrendous new UI on me, and the "Classic UI" plugin I have to use to keep my IDE working the way I have been used to for the past 15 years is never updated at the same time.
ThePhysicist 22 hours ago [-]
There's this great blog post by Joel Spolsky from 2000 [1], where he essentially argues that controlling your environment makes you happy. He writes about his summer job in a bakery and how the dough mixers would be so unpredictable and how frustrating that was. I think AI agents are quite similar to a lot of folks, they change significantly with each major model update and even every day as the vendor tweaks the system prompts and settings, so you never feel "in control", it's more like pushing buttons on some blackbox and hoping the right stuff happens inside. Most people are unhappy about that as it takes away the mastery and craft aspect of software development and makes them managers of unpredictable AI tools. I certainly get this feeling even though I like AI in general, but having days where everything goes so well working with the agent and then days where nothing really seems to work and not knowing why is quite frustrating.
> and even every day as the vendor tweaks the system prompts
I just patched Claude Code's system prompts, pinned the version and stopped upgrading without first dissecting and auditing the executable. Even discovered Anthropic can remotely inject strings into the system prompt via some "growth book" or something. Neutralized that too.
Things got a lot better after I started doing this. It straight up fixed Opus 4.6, and Opus 4.8 got more consistent in my subjective experience.
Sadly there's nothing I can do about Anthropic's server side "system reminders" whenever some prompt trips their classifiers or whatever.
I'm in the process of switching to OpenAI and Codex. The open source harness is a breath of fresh air. We'll see how that goes.
GZGavinZhao 20 hours ago [-]
Maybe you should try using Pi [0] if you care a lot about owning your harness and making sure you're in control of what goes into your system prompt and context?
I don't think you got what the parent comment was saying.
slopinthebag 20 hours ago [-]
Crazy amount of effort when you can just use a different harness and a better and cheaper model.
matheusmoreira 20 hours ago [-]
"Cheaper" compared to API prices, right? I've run the numbers and the frontier subscriptions are still the best option. 100% usage every week is a truly absurd amount of value. Unfortunately the subscriptions can't be used outside the official harnesses.
kelnos 17 hours ago [-]
A month or so ago I asked Claude about using open-weight models through another provider (after describing my usage patterns), and, amusingly, it told me that I could probably save money ditching my Claude Max 5x sub and switching to something like Fireworks, using GLM 5.2, even if I'd be paying API rates.
I still haven't gotten around to trying it, though, so I don't know what the reality is. And new open-weight models have been released since then that I'd want to evaluate...
zahlman 12 hours ago [-]
Trying to force myself to find ways to extract full value out of such a subscription sounds very unpleasant.
matheusmoreira 9 hours ago [-]
It's the literal software development gacha, habit forming, timed intermittent reward schedules and all. Making insane progress on my side projects though.
If I have too much usage I'll simply find something for the AI to do. Guided router update. Router hardening. Laptop hardening. Penetration test my router now that I somehow got into the cyber program. Local models research. PC build research. Smartphone research. Financial investments research. Business research. Medical research. Laptop firmware reverse engineering. Old video game reverse engineering. Pick a random open source project and let's explore the code base. Analyze all of my ten thousand HN posts and tell me interesting facts.
If I can't think of anything I default to picking a random git repository and launching a massive parallel code review session. That's guaranteed to kill any remaining usage in exchange for useful output. Then I can go enjoy my weekend guilt free. Unless they reset the usage.
slopinthebag 16 hours ago [-]
Both API and subscription prices are cheaper with all other models and providers compared to Anthropic.
GPT is also better than Claude from my personal experience. GLM 5.2 is equal and Kimi K3 is better.
But I mostly use Deepseek and pay api prices.
zahlman 12 hours ago [-]
This is, I think, a big part of why I wouldn't want to remove myself from the loop.
firasd 24 hours ago [-]
I feel like these abstractions like "CI might not work well in the era of agentic tooling" are fine for thought-leadership posts but there's so much hands-on work to be done. The last word on AI computer use shouldn't be bash utils that were already feature-complete before MJ recorded Thriller.
There is some movement in this direction--there is a new 'gh' subcommand called repo read-file for example, that lets agents view a file without cloning a repo. And I made something called venetianblinds that shows equidistant samples of a file. In combo they work pretty well:
--- sample 2/20 char 283427 line 5855 col 53 range 283367:283487
le].
**
** ^Closing a BLOB shall cause the current transaction to commit
** if there are no other BLOBs, no pending prep
^
bob1029 22 hours ago [-]
Hand-crafted, domain-specific tools massively outperform general purpose ones.
Shell execution and raw DOM access are great for a backstop, but you can go so much further with just a little bit of translation and delegation around the environment.
I think browser automation is probably the most apt scenario. Often a human who understands how a page is meant to be perceived can transform a megabyte of raw web content down to a few hundred bytes of plaintext without any reduction in fidelity. This can be achieved using deterministic code that is guaranteed to provide a perfect transform every time.
The performance difference between raw DOM access and curated plaintext is like a step function. With raw access you get maybe 10-15 steps into a complex workflow before the wheels pop off. With curated access I've seen it go 100+ screens without issues.
Simply managing the token bloat is probably the most important objective here. If that's all you focus on it will probably go really well.
firasd 20 hours ago [-]
Yes exactly. It's almost like there's this expectation that if the LLM GPU-Maxxes enough it can do anything but like... why. There's no reason to wade through a million HTML tags if you can slice out the exact div you want to check
p1necone 22 hours ago [-]
I've found the process of building good reliable CI that thoroughly covers everything has been greatly improved by LLMs. There's so much tedious plumbing and grunt work involved in building CI and automated testing infrastructure for bespoke products that they can handle just fine while you concentrate on the important bits - I would say it's an area where agentic workflows are even more suited than regular product code.
LtWorf 5 hours ago [-]
100% coverage and useful tests aren't the same thing. Why are we still not aware of this?
hahahaa 24 hours ago [-]
MCP is the answer to not using bash, right?
Bash is a great control surface anyway for LLMs as it is wordy and powerful.
firasd 23 hours ago [-]
Yeah as far as what I'm talking about bash CLI vs MCP doesn't matter (there could be a file_sample MCP tool)---I'm saying that by default Windows, Linux etc don't have this venetianblinds affordance of seeing equally spaced samples. It's a trivial algo but it's very handy these days cause LLMs can't really just be like 'okay I'm gonna open this file at random and scroll around'--the file itself is an unknown blob (JSON data, Python, Typescript, a log file etc) without coordinates. So they fall back to thinking they've gotten a good sense of the file from head/grep or they write ad-hoc Python to manipulate the file.
Another venetianblinds survey, of the Paul Graham 'What I Worked On' article that's used in many LlamaIndex examples:
npx github:firasd/venetianblinds pgworkedon.txt
--- sample 1/20 char 0 line 1 col 1 range 0:60
Before college the two main things I worked on, outside of s
--- sample 2/20 char 3946 line 19 col 237 range 3886:4006
d an intelligent computer called Mike, and a PBS documentary that showed Terry Winograd using SHRDLU. I haven't tried re
LelouBil 23 hours ago [-]
Also, this is a trivial example but I would rather an LLm call a recursive "grep" than using 1000 read_file MCP calls.
Having a scripting language as a tool is powerfull, and can help remove unnecessary stuff from the LLM's context.
ElectricalUnion 23 hours ago [-]
Problem is that bash is too sharp to handle to smart and gullible clankers without a sandbox - That I think everyone should be using anyways, for everything, even things not related to clankers - Android and Qubes are right. The app/vm, and whatever it tries, should not be considered trusted by default.
applfanboysbgon 23 hours ago [-]
Qubes is directionally correct, Android is extremely not. Safety must not be obtained by preventing users from controlling their own computing devices, or else we face a dire future.
inigyou 23 hours ago [-]
This is a lot of words to say absolutely nothing. Seems apropos for Stack Overflow though.
groundzeros2015 21 hours ago [-]
When I see “stack overflow blog” the brand encodes distrust. I remember when they removed links to meta and replaced them with the corporate blog with low quality PR and political agitation.
TacticalCoder 18 hours ago [-]
> When I see “stack overflow blog” the brand encodes distrust.
Incredible, in a way, that you still see it as a brand. I honestly thought it was a 10 years old blog post somehow making it to frontpage (as some old blog posts sometimes do on HN).
Who still uses DeadOverflow-full-of-outdated-answers? (outdated and often just plain wrong too)
I haven't been there in like 15 years or something. Feels like ages.
It's owned by some really stupid private equity firm since about 2019, which is trying to revive it as a platform for agents to talk to each other.
utopiah 1 days ago [-]
Im typing this in GVim thanks to Tridactyl using my new mechanical keyboard running a ZMK firmware I just built via Github actions (or directly via ZMK Studio).
This is ridiculously complex to just type a few paragraphs. Nobody in their right mind would invest this amount of yak shaving... and yet I do so because I bet, rather confidently, that in few years, heck few decades, all those tools will be different (or maybe not, I still use Vim on my server, desktop but even mobile phone) but the lessons will remain practical.
IMHO the trust comes from trust yes but also more directly plain ownership.
zahlman 12 hours ago [-]
> typing this in GVim thanks to Tridactyl
Thanks for the heads-up that such a thing is possible. I will definitely be investigating it.
utopiah 9 hours ago [-]
With pleasure, ^i in a textarea or even input element brings your text in GVim then saving brings you back. Really handy feature to bring regexes or text navigation to any Webpage.
alexpotato 20 hours ago [-]
There is a great article called "Manual Work is a Bug: Always be Automating" [0] that was written in the pre LLM era for technical operations teams. I would argue that it is just as relevant today as it was then.
To summarize:
- start making a list of the manual tasks you do
- if those tasks involve running command line tools, add an item with the commands you run
- if they are manual tasks, add those too
- over time, keep automating one portion of the list at a time e.g. the commands can become a script, the manual tasks can become tickets to another team to automate etc
At the end of the above process you have a series of automated steps that become a system instead of a bunch of items in someone's head.
In my mind, the only thing that changed with LLMs is that it's faster to create the scripts and some of the manual tasks can be done by the LLM until you get a script to do that too.
We invented code to help do "mechanical" tasks over and over again in the same way. Why replace that with agentic systems??
P.S. This is also why the whole "just commit the prompt, bro" is such utter hogwash
Like all advice, A.B.A. some pretty serious flaws if you apply it too universally. Automated solutions tend not to replacing manual work perfectly or completely. They do most of the same things, and the distinction tends to be forgotten. Maybe a human goes in and fixes things for a bit afterward, but that doesn't last. Eventually the accessible 90% is declared good enough and people forget that the rest is even possible.
The first case of large scale automation demonstrates this very well. Medieval manuscripts looked like this [0]. You can see the imperfections, but it's still a beautiful book. Gutenberg bibles omitted the scribe in favor of automated printing, but humans remained involved with the creative details in illumination and rubrication. The result is genuinely beautiful [1]. The works that were printed a century later rarely featured this kind of post-print human involvement [2]. This is still better than most of what's printed today, but it's a clear step down.
Now, there's a reasonable argument to be made that the quality differences don't matter for books and certainly don't outweigh the cost advantages. But imagine someone at your passport office has taken A.B.A. to heart and automated 90% of the job. Over time, the organization will stop handling all the edge cases and. But the edge cases didn't stop happening, they're just no longer visible. Beyond a certain scale (e.g. Google and other FAANGs), those inevitable failures manifest as seemingly capricious behavior that makes everyone hate your system and produces outcomes no human involved actually wants.
As a teacher I know for a fact some people just learn faster than others. So I think part of why developers get attached to tools is because learning takes time. Some people can pick up new tools in a day. I can't. I feel like new tools are created faster than I can learn them.
win311fwg 18 hours ago [-]
All tools require time to learn, but not all tools gain attachment, even when one has committed to learning the tool. From my observation, attachment is formed when one finds something about a tool that feels like a secret insight or advantage that is being overlooked by most others, thus establishing a tribe of those who have "seen the light", with the tool becoming the identity of the tribe.
Mikhail_Edoshin 7 hours ago [-]
Developers can also build tools as needed. This is the key to simplicity.
oooyay 23 hours ago [-]
I'm not sure the difference changes the conclusion but I think projects demanded certain workflows and resultant processes, not the other way around. That's why Jetbrains has so many workstream specific IDEs that sold very well. The processes didn't go away but a lot of us changed our IDE surface. Those processes still need to exist, largely, but the way in which we invoke them is moving and changing. To some degree, the processes are also changing because other factors are changing outside of the tooling.
For example, I use Codex and Claude Code by default, but when I need to look at the API surface, read tests, etc I have those tools setup to open Zed. Zed is also rapidly evolving in the other direction, where it's closer to the tools that are opening it. It won't be long, I think, until I can continue my prompt from inside Zed.
zahlman 12 hours ago [-]
Stack Overflow was such a tool, prior to being bought by Prosus and dealing with Monica so absurdly.
Trust lost is rarely ever regained. That they publish a piece like this on the Stack Overflow blog is insult to injury.
1saadcodes 19 hours ago [-]
I think the article is right that developers become attached to tools because of trust. Ironically though, that's also why many people drifted away from Stack Overflow. The answers were trustworthy, but over time the experience of asking questions felt less and less welcoming than searching for existing ones
kristianc 22 hours ago [-]
Stack Overflow conveniently defining the problem in a way that makes SO the answer. The irony is that the kind of code review that LLMs make possible would never have been doable on SO lest you be accused of trying to 'outsource the work'.
Now agentic coding produces the full 100 line change, and SO says the real crisis is that nobody has reviewed it properly. SO was fully responsible for creating the culture where no one wanted to use it.
22 hours ago [-]
iwontberude 4 hours ago [-]
Tools encode determinism, not trust
pjio 14 hours ago [-]
> The good news is that work gets done faster. The bad news is that we don’t trust it when it’s done.
It's not done when you can't trust it.
niemenghui 16 hours ago [-]
I agree: With AI, we’re able to automate a lot of those processes. The good news is that work gets done faster. The bad news is that we don’t trust it when it’s done.
Avery29 12 hours ago [-]
Trust in tools often comes from predictability and shared context
willjp 22 hours ago [-]
I agree with this article, but it's so much more than just trust in a workflow.
It's an API you can trust, in an ocean of change.
That's what you're choosing; an api, abi... whatever you call it.
You're committing to an extensible boundary.
And both the boundary, and the extensibility (not to mention the openness) are why they win.
And shelling out to a cli? That's a pretty damn compelling and composable boundary, IMO.
shostack 23 hours ago [-]
An example of this in action is the utter inability to get deepseek v4 flash (even the new version) to stay concise. I have jumped through all sorts of hoops with deterministic checks, pre-message injection hooks, memory framework, etc and when it fails still and I ask why it essentially says "I forgot."
This makes it unreliable and preferences are things I need to assume are treated as exactly that, preferences, not hard settings.
It is an area where it is more like working with an unreliable human than I would prefer.
mlloyd 22 hours ago [-]
Lots of people saying the same about Deepseek v4 Flash. Seems like it's an artifact of that model.
overgard 1 days ago [-]
I have to admit, I asked ChatGPT to do a TLDR summary because I found the writing meandered quite a bit. I think the overall point is sound:
> "Developers become attached to tools like Vim, Emacs, or an IDE because years of experience make those tools predictable extensions of their thinking. The attachment is less about features and more about accumulated trust, muscle memory, and a workflow built around known boundaries.
> AI coding agents disrupt that trust because they are fast but probabilistic, opaque, constantly changing, and capable of producing more code than humans can realistically review. This shifts the bottleneck from writing code to specifying, reviewing, validating, and operating it safely."
(Note the > is paraphrasing)
Trust is a big problem I'm having with these tools so far. What I've been running into a lot is, I'll get the equivalent 40 hours of work done in 8 hours, and I'm like, wow, that really was quick. Then I'll start using the application I'm making more directly (a tool for writing), and I'll start to see that it's broken all over the place in very surprising ways (ie, updating this menu item broke something on the other side of the app, etc.). So then I spend another 40 hours of real wall clock time kind of fixing everything that was broken, and at the end of those two weeks I'm like, did I actually go much faster or was that all kind of a wash? Because if I'm not going faster in overall terms, then the loss of deep understanding of the code base might not be worth it if my pace is the same.
I'm sure someone is going to be like "BRUH AUTOMATED TESTS" or "BRUH MODEL CHOICE". I have a LOT of automated tests, and I don't like fussing with models so I pretty much use Opus on high reasoning for most things (or the equivalent from other providers). Code review also doesn't help that much, for much of the same reason it doesn't tend to help find bugs in human written code either.. you're reading the happy path usually.
Anyway I wouldn't say these tools aren't useful, but, I'm deeply skeptical of all the productivity claims because I think people just look at one dimension of it while ignoring all the other important dimensions. Yeah you can generate a lot of crap fast, but most of it is not shippable and making it shippable does take time.
derek1800 1 days ago [-]
If you are spending 40 hours fixing everything that was broken, the question I have is does your AI tools have the necessary context to be successful and not result in a lot of broken items?
Also, is there ways for AI to help prevent the loss of deep understanding of your code base without you having to know every line of code deeply?
overgard 23 hours ago [-]
I was a bit hazy on numbers because I don't scientifically record them, but I guess what I'll say is the fixing and verifying takes a lot longer than writing the initial code. This was true before LLMs, but when writing code by hand I had the context of the code in my head, so potential issues, blast radius, etc. was a lot more obvious. By definition most of the things that break are things that are not trivially testable. Unfortunately, it's not as easy as saying "Claude, the thing broke, plz fix"; I've had to spend a lot of time recently helping it with context from debuggers, or just debugging myself manually, adding log statements, etc.
Am I giving the agent enough context? Well, I'm giving it as much as I can. Each submodule has an AGENTS.md, I have the agents add gotchas and instructions for some feature work when I discover where an agent went wrong, the codebase has a lot of comments along the lines of "if you edit this section, you need to also edit XYZ", and I lean on the type system as much as I can to make wrong-code not compile. It has access to playwright for driving the UI if it wants to. (Weirdly, I've found that Claude is really inconsistent about using these tools -- even though the instructions make it clear that it's allowed and encouraged. I think if your workflow differs from the models training and thusly you have to tell it so in AGENTS.md/CLAUDE.md, then it's very inconsistent about following those instructions. For instance, I don't want Claude to commit and I don't want it to sign commit messages, and it still does that all the time even though it's my like #1 directive of "don't mess with my git history")
There are some things though that are very hard for it to test. I'm exporting essentially a programming language to three game engine runtimes. They all have automated tests, but, I think people that have worked in video games know that games are very hard to automate testing on. This isn't really the fault of the agent I would say, just the nature of the problem, but it is worth noting.
I guess this is a long winded way of saying, even with LLMs tech debt is a thing you have to manage, and I think managing tech debt becomes even more important when you're dealing with LLMs, not less important.
dijit 1 days ago [-]
"40 hours" in his context here is actually a work day, so 7-8hrs.
He says "40 hours" because he feels like he's managed to do 40 hours worth of work in this time, but then has to spend another "40 hours" (actually: 1 day) just going around kicking tyres.
Obviously the implication is that it's a net gain of some kind, but he's unsure if he caught everything.
(sorry to reiterate the GP, but I feel like you missed the important nuance that it's not a real 40 hours of time).
grey-area 24 hours ago [-]
No the second 40 hours is a real 40 hours (two weeks), and the implication is there is no real time saving.
inigyou 23 hours ago [-]
Wow, just wow. This is the first time I've encountered this particularly AI apologism. To recap:
Alice: "in the end, AI doesn't make me any faster because it still takes 80 hours to do 80 hours of work once I fix it"
Bob (AI booster): "actually you might've been holding it wrong, did you try XYZ?"
Carol (super AI booster): "Bob, actually Alice means it took 16 hours to do 80 hours of work. So it did work for her."
Alice: "no I fucking didn't"
overgard 23 hours ago [-]
Sorry, my original phrasing was confusing which you should not be downvoted for. I've edited my original comment to clarify what I meant (hopefully).
zahlman 12 hours ago [-]
> I'll get the equivalent 40 hours of work done in 8 hours.... Then I'll start using the application... and I'll start to see that it's broken all over the place in very surprising ways
I would say that this means that it only appeared that you got 40 hours of work done.
> I'm sure someone is going to be like "BRUH AUTOMATED TESTS" or "BRUH MODEL CHOICE". I have a LOT of automated tests, and I don't like fussing with models so I pretty much use Opus on high reasoning for most things (or the equivalent from other providers).
Bruh, manual tests. It shouldn't be an entire day until you "start using the application more directly". It's rare that you really know what you want until you try something that isn't what you want, and figure out what's wrong with it. LLMs aren't privy to your thought patterns, so they can't help with that.
pmichaud 1 days ago [-]
I was surprised to see you say you have automated tests. To me this makes most of the difference, but you have you actually have a good test suite, like one that actually proves the code does what you want it to do. Unit tests, property tests, e2e tests. The other part that makes all difference, is you have to be all up in the model's business about architecture. Pick something that wants to testable and isolation friendly, data models that are correct by construction (ie invalid states are not expressible), etc. It absolutely will try to cut corners give you bullshit slop at every turn, you have to keep the structure sane and build the right tests and harness around it.
And I can hear your objection now: correct, it's probably not worth all that for a throwaway, but the effort per output goes down as the infra builds up and you end up with a program that can reliably expand.
overgard 23 hours ago [-]
I think testing is really important (even without LLMs). Currently I have 2247 unit tests across 132 files in a ~150K LOC codebase (test code included in that count). It takes about 50s to run. There could be more, but it's not nothing. There is playwright for it to test drive the UI, although if I'm being perfectly honest even before LLMs I thought that kind of test tends to be brittle and annoying to write (I guess I don't have to write them anymore, but they are still brittle). I'm honestly trying to give it as much structure as I possibly can -- I'm not trying to setup the agent to fail so I can be like "gotcha!"
I'll also just point out my philosophy for using LLMs for this project, which is that I'm not trying to go as fast as I can. (I want to go at a good pace, but this isn't an experiment to just finish something over a weekend). The 150k LOC have come about since February, with some mix of me writing code and LLMs, so on average I'm probably bringing in about 800 LOC per day, which I imagine a lot of vibers would find to be glacial. To me that's the sustainable rate of what I can do when you factor in that I need to test drive every feature, make sure it doesn't conflict with another feature, check for bugs, check that the code looks reasonable, and debugging. (I also think that rate limit is specific to this project: I could see easier to test things going much faster, and harder to test things going slower)
pianopatrick 22 hours ago [-]
Isn't being brittle kinda the point of playwright tests? Because each test uses the whole app those tests can catch errors that happen at many places in the code. But that means lots of things can cause a test failure.
tommyage 22 hours ago [-]
Wait; So you say you outsource development to a LLM fully knowing that it will not write sound/deterministic code and your quality requirement are your test cases, right?
I further suppose you are not writing the test cases in its whole by hand. But you try to specify them before you hand out the development task to the LLM, right?
It sounds like you just introduces a new team member which is not trustworthy yet.
Normally you would review each change of him and explain how to improve hisself and the code.
But that's not possible to a fixed-state LLM.
That sounds exhausting.
If your Markdown Files should aim at improving the LLM contributions, you are again stuck with the fact that it did not follow in the first place.
So I conclude: You pay for an Intern who is not trustworthy and pay additional input token on Markdown Files to still no be certain about future contributions.
I just don't grasp how we as engineers are accepting this and integrate it into our craft.
And: Everybody using External LLM Services, possibly providing the entire project as context, is allowed to let the source code of your company be leaked to some third party. I hope that party is trustworthy and does not have a track record of copyright infridgement. Because this would be fairly naive and reason to be fired. So we as Engineers knowing the implications should therefore point to these issues at the correct management level.
At least that's what I am doing shrugs
mlloyd 22 hours ago [-]
I have it give me a punchlist and a tldr in a table format for every turn. It helps cut through the incessant babbling.
skydhash 23 hours ago [-]
> Anyway I wouldn't say these tools aren't useful, but, I'm deeply skeptical of all the productivity claims because I think people just look at one dimension of it while ignoring all the other important dimensions. Yeah you can generate a lot of crap fast, but most of it is not shippable and making it shippable does take time.
My own stance is that there's never any reason to go fast on anything. Communication has always been the bottleneck. Whether it's about gathering requirements or understanding the purpose of a badly written code, any speed improvements I get has always been a small percentage of the overall progress.
What has helped more is my understanding of the platform and some theoretical knowledge. Because one I get the information, I can quickly derive a solution in my mind. And that solution has always been easy and fast to implement, at least the happy path. 90% of the time taken in coding is always about handling all the edge cases, aka fixing bugs. And writing tests so that you're not easily introducing more bugs.
nvgjbdhmkdd 1 days ago [-]
[dead]
alyx 1 days ago [-]
[flagged]
kittikitti 24 hours ago [-]
I operate on zero trust because I find that people won't trust me regardless of what their stated reasons are. They just feel uncomfortable. On top of that, people will hallucinate things in order to not trust me.
I update my toolset all the time. It always results in discomfort and backlash but people don't understand that the goal isn't their perception or trust. It's about skill, ability, and execution. This idea probably won't get me promoted but it will get me paid. I am not attached to tools because I learned the hard way that they will always find a way to take them from me.
Zero trust is a better alternative for people like me. In terms of cybersecurity, being attached to a tool is crutch because fatal flaws in every design are frequently found. As it relates to agentic AI, I never select the "Yes, trust the AI and let Claude execute arbitrary commands in a non-sandboxed environment" option. However, I frequently utilize agents, but I'm not going to have the "Jesus, take the wheel" moment with them right now. That being said, AI is a very helpful tool that helps me create boilerplate code, brainstorm ideas, and review my work. I also anticipate when AI can, in fact, take the wheel and I'm looking forward to it.
Parallel to this, I also know that developers often disagree, and I'm not casting judgement on anyone for being attached. If it's Turing-complete, then I have the background to complete the task. In these scenarios, I just adopt whatever tools work best in team building because, in my own words, I'm not too attached to the way I do things.
hahahaa 23 hours ago [-]
What constitutes taking the wheel? Skip permissions in a proper sandbox is fine IMO. There is a small amount of risk I admit though.
antonvs 17 hours ago [-]
It’s hard to take an article seriously that starts out admitting an earlier troll. This is clearly a person that’s more interested in engagement than content.
sharpnick 6 hours ago [-]
[flagged]
16 hours ago [-]
tizerluo 17 hours ago [-]
[flagged]
p1necone 22 hours ago [-]
[dead]
fibuladev 23 hours ago [-]
[dead]
acchow 20 hours ago [-]
[dead]
hamza7159 24 hours ago [-]
[dead]
youareinsuffera 24 hours ago [-]
[flagged]
zephen 23 hours ago [-]
> Apparently you just deserve a life of pain.
Don't we all?
(I can see that you are starting to get downvoted as well. Spread the love.)
Stack overflow was interesting. Its design was the only thing like it at the time, and made it a Schelling point for programming knowledge distribution, but also a welcoming environment for the programming equivalent of grammar nazis.
Some of those programming nazis, of course, had suffered at the hands of previous ones on stack overflow before becoming "enlightened." And thus, the generational hazing began.
It was great if google directed you to exactly the right answer, but god help you if you couldn't figure it out, and posed a question that someone thought didn't contain an MCVE.
Also, a few too many of the high-reputation people would post complete garbage on topics they knew absolutely nothing about.
inigyou 23 hours ago [-]
You can't just call everyone you don't like a Nazi.
zephen 22 hours ago [-]
The term "xxx nazi" has been around for over 70 years to describe people with sticks shoved so far up their asses that they poke out the tops of their heads.
And "grammar nazi" has been around since at least 1990, and the infamous Seinfeld Soup Nazi since 1995.
I'm not sure of the etymology of the "not everybody's a Nazi Nazi" but I'm sure you're not the first.
But in any case:
> You can't just call everyone you don't like a Nazi.
Yes, yes, I can. I don't (because I reserve the appellation for certain particular kinds of attitudes), but I could if I wanted to.
fitsumbelay 1 days ago [-]
[flagged]
inigyou 23 hours ago [-]
Yes that is what happened to SO. They now get practically zero questions, zero answers, zero page views, and zero ad revenue. It is their own fault because they froze everyone out of the site and then once LLMs became an alternative, everyone started asking their questions to LLMs. They are now trying to somehow pivot to AI to make revenue again, starting with Stack Overflow for Agents, and now with wordy vacuous blog posts to show off how AI they are.
>If your kitchen knife kept changing shape, weight, and edge, you’d have to relearn it every time; that’s a hard tool to build trust in.
This concept was betrayed far before agentic tools, with a much earlier concept: Automatic updates.
To use one product as an example: When Windows ME and Windows Vista came out, people hated them even more than they usually hated Windows, so they did not use them. Microsoft was forced to respond by making a not-quite-as-bad OS in Windows XP and a pretty good OS in Windows 7 respectively. No longer is that an option, your workflow will simply be interrupted by automatic updates.
>Vim and Emacs, in their infinite customizability, can be molded to fit your exact hand and workflow
Vim is one major exception to the automatic update problem. I trust vim not just because it can do a ton of shit (although that is certainly nice), but because unlike most other software, its UI doesn't change unless I tell it to change. Aside from switching from vim to neovim (my decision, not a forced update), my muscle memory from a couple decades ago still works today.
It is "velocity fallacy" — product people want "all the features ASAP or right away".
Until users with their managers come with pitchforks and torches. I worked on such internal project where we as developers were able to deliver new features and new version every 2 weeks (which is not a pinnacle of the game of course) and were thinking if we can move to daily delivery. Because we were good devs and wanted to appease product owner.
Until one day product owner came back with feedback, how everyone is pissed off at him for shifting ground under people feet, while he also thought he is doing great delivering all those great features. It was pushed back to limited amount of features each month.
People need training, people need to understand what and why.
In the end it turns out it is also AI coding fallacy, because most of the software is built for limited audience, which has its specific timeline on accepting and internalising new features or different ways of doing stuff. Unless we take humans from the equation and we start building for AI itself.
We published 1-2 releases of our component each month. Eventually another (internal) team would pick it up along with others, test, and release the combined set maybe 1-2 times a year.
Customers being extremely risk adverse never wanted to update, due to risk of changes combined with the interruption, even though we were fixing serious bugs left and right from the earlier rushed development.
We’d get escalations on things that were fixed years ago.
Probably should have spent another 3-5 years getting the first release in better shape instead of spending 3-5 years flinging flaming turds to the paying customers.
The executives slowly released they inexplicably lost half the business compared to the previous generation as upstart competitors stole their market share.
You know that famous “how shit happens” tale? I’m pretty sure each layer of management was telling the next how powerful the product was… few could abide by it.
> Probably should have spent another 3-5 years getting the first release in better shape instead of spending 3-5 years flinging flaming turds to the paying customers.
I wonder if the company would have had the money to pay salaries for 3 years, unless for those "paying customers" that obviously started paying much earlier than that.
And
> The executives slowly released they inexplicably lost half the business compared to the previous generation as upstart competitors stole their market share.
s/released/realized/ ? Exactly one of my worries. Maybe they could have created those competitors themselves. Two brands, one for the original product and one for a product with a somewhat different layout and all the improvements that they did not dare to add the the original product. But then you need a third brand, a fourth one, etc.
We're still at the "search" phase, there is very little wisdom at social and personal levels.
It's clear from the rapid flux in tooling, methods, ideology, and the results, that we basically have no idea what we're doing re: using computers and building software. A lot of the self assuredness of current advice is self soothing behavior.
This is so weird to me, though. Your customers had two choices:
1. Stay with your software and upgrade to the latest version, where the bugs they were hitting were fixed, and risk some amount of retraining due to UI/UX changes.
2. Switch to a completely new, different product, which guarantees retraining (possibly a lot more), and probably would contain the same or different bugs they were fighting with when using your software.
And... they went with #2?? I get that people think the grass is always greener on the other side, but there's a reason why we have that saying!
As anecdote of one, I was once part of a project to port a .NET Framework to Java, because the customer was really annoyed with the rewrite, as the application relied heavily on the .NET Features that never made the cut to modern .NET.
Another two .NET heavy weights in .NET CMS space, Sitecore and Optimizely, nowadays rely on JS/TS frameworks for their extensibility SDKs on the SaaS products for headless deployment, only the classical (older) PaaS still support .NET as extension language.
Modern software is quicksand. Maybe an update improves things, but users have all experienced working software made worse.
And then new features are released as MVPs - meaning, they do the absolute minimum to check the box (almost a literal box: close the feature ticket internally), and inefficiently so. Whether it'll be iterated on afterwards, depends. Pretty unlikely in the immediate term. They need to "collect usage data to know what to do" first, which means it gets deprioritized relative to moar features.
But that's fine for me. As long as copy-paste works, I can make do with other software, perhaps competing software, or worst case, have Claude find and use some powerful-but-janky OSS CLI tool to do the stuff for me.
I'm less angry at it than I used to. Mostly because I don't have time to be annoyed anymore, but it's true that I've learned to like updates from few companies. That's because after years - years - I've noticed things gradually improving on average. True of Android & Samsung OneUI, except when it's not. True of UniFi stuff. If I see some new feature behaving badly, I now mostly trust they'll eventually fix it. It'll take a year or three, they'll overhaul it entirely twice, and I may need to buy a new phone to get it, but it will happen one day. But most software doesn't even clear that bar.
Someone whose partner is about to leave them loudly proclaims "I've changed!". Should they believe them? Often the answer turns out to be "no".
It's quite difficult to actually build back trust in these situations.
I literally was told not to narrow down the features too much or the product would be too constrained... well then wtf does it actually _do_ for anyone?
The idea is that the user should be able to be productive more or less immediately after installing the software, and then incrementally customize it as they discover they have different needs than the defaults give them.
A lot of developers/companies forget the first part, and ship something that requires a team of consultants six months to set up before anyone can do anything useful with it. Of course no one wants that. (Well, except for the highly-paid consultants.)
This is a classic IDE problem. The IDE has a preferred directory layout it wants, and some other build tool you're using has a different directory layout.
The problem seems to be that there's not many such great abstractions to be found.
I have seen dozens upon dozens of Excel sheets which "just worked" until they didn't and then of course person who could fix that left company 10 years ago.
Besides I also know dozens of people whose life would be easier if they would learn a bit more of Excel like pivot table is there for them.
So which is preferable - customisability or very limited feature set?
There is no general answer that will answer which is preferable.
It is easy for people to come up with features they want or a customisation and they don't care for the cost of actually maintaining that feature or customisation.
I work on building SaaS platform, we had multiple customers for whom we build custom features and they paid for development of those features. Not fun part is after 2-3 years those customers are gone (for example employees at customer rotated and they switched to something totally different) — but now we are left with dead feature no one is paying for supporting, some are easy to remove, some are not.
But I would not expect a large percentage of developers to want to do that. I've used IDEs as well, and they're mostly fine, and work really well for lots of people. Many IDEs are pretty damn customizeable too, though most people may not change many settings.
So I think there's a place for automatic updates, but the feature should be restricted to security fixes only. Using it to foist ill-conceived changes on users is just abuse.
My late dad preferred to not upgrade at any cost to the end!
Now today of course things are quite different and it’s not unusual for a routine Windows update to turn things upside down. Users are more justified in update-averseness than they were 15-20 years ago, except of course now that brings much greater risk of getting pwned than it did back then.
I don't think "ill-conceived" is giving enough credit. It has been clear to me for a long time that "security" is the justification for pushing the user to update, but the updates themselves are frequently leveraged as a vector for other, less user-friendly, practices.
I think we would be better off if people worked harder to prevent vulnerabilities before releasing the software.
Pawn Stars meme: "Best I can do is people using glorified chatbots to generate mediocre code an order of magnitude faster"
That said, the vast majority of the exploits that led to the widespread adoption of automatic update mechanisms were based around memory safety bugs which we do in fact have solutions to entirely prevent in most newly developed software these days.
Prior to XP, MS had two lines of Windows: the Windows 9x kernels and the Windows NT kernels. Windows XP was meant to be the merger of the two lines, adapting Windows NT to have compatibility with Windows 95 and Windows 98 features. Unfortunately, Windows XP development went overlong, so MS wedged in Windows ME to give a stop-gap release until XP could actually be released.
This timeline doesn't seem to make sense, as Windows XP came out in 2001, and Windows Vista in 2006-2007. Maybe you are referring to Service Pack 3 in 2008?
> Microsoft discussed a plan to delay Neptune in favor of an interim OS known as "Asteroid", which would have been an update to Windows 2000 (Windows NT 5.0), and have a consumer-oriented version. At the WinHEC conference on April 7, 1999, Steve Ballmer announced an updated version of Windows 98 known as Windows Millennium, breaking a promise made by Microsoft CEO Bill Gates in 1998 that Windows 98 would be the final consumer-oriented version of Windows to use the MS-DOS architecture. [0]
So XP was not a reaction to Me's reception, it was already in the works as a replacement when Me came out.
[0] https://en.wikipedia.org/wiki/Development_of_Windows_XP
I did try XP here and there, but was instantly turned off by the cartoonish default theme (yes, I know you could change it). It was buggier than Win2k, and I didn't have the patience to wait around until they fixed it. I was told later on by people in the know that Service Pack 3 was the bees knees, but by then I'd moved on to Linux full-time (with some jaunts into OS X) and felt no need to come back.
> When Windows ME and Windows Vista came out [...]. Microsoft was forced to respond by making [...] Windows XP and [...] Windows 7 respectively.
It's basically from ME and Vista to XP and 7, respectively. AFAIK respectively in this context means that for ME, they were forced to respond with XP, and for Vista, they were forced to respond with 7.
Also to respond generally to other posts. I am aware of the the separation between NT and 9x. That was Microsoft's problem and not the customers' problem. They could not force people onto ME and had to actually appeal to customers with XP. Then later on, they could not force people onto Vista and had to actually appeal to customers with 7.
Automatic updates remove the step where they have to appeal to anybody.
And I don't mean to single out MS. I remember having a Mac and switching from some version of OS9 back to 8.6 for some reason (don't recall why, but it doesn't matter because it was my computer so it was my decision). Nowadays people are complaining about Liquid Glass and they cannot rollback their OS on their iDevices.
UI changes were pretty minor, especially compared to the flat design of windows 8.
Except for one thing that pissed me off a great deal. I am not a true vim user, though I use it from time to time, because it is everywhere and it works through ssh. At some point they introduced "visual mode", and it turns on when you start using the mouse, it broke the way I used the mouse copy-paste in vim. I know I can do "set mouse-=a", but when I am just using vim as "the standard text editor" (sorry ed), I don't want to configure anything as it is usually a one shot job.
I never asked for that, at some time, it just happened. I guess as a major upgrade, but the thing is, something changed that I didn't want to change.
I understand the reason for this change, also https://xkcd.com/1172/ but I just wanted to say that even vim is not immune.
Some tools are immune though. Usually command line tools used in scripts. For example "apt-get" doesn't change, but "apt" does. "apt-get" is what you use when you want a stable interface (especially scripts), "apt" is for when you want something nicer.
When I first started at my company, we did all work on shared VMs, and the system vimrc had that "mouse" setting enabled. Something an employee had done decades ago to be helpful, really confused me until I realized what was going on. I'm thinking your distro, not vim, is what flipped the setting.
(For bonus points, you can make a program to copy ~/.ssh/config around as well so that your config gets copied to servers you ssh to from there.)
My entire reason for switching from vim to neovim was that vim did change its UI, by neutering /etc/vimrc in a major update a while ago.
The thing about it, though, is that LLMs aren't "your" tools. They're someone else's tools, and you are at the whims of day-to-day changes to them.
Vim and emacs aren't unique in this. In fact quite a few non-commercial FOSS projects don't make significant changes to the UI. I think there are (at least) a couple of reasons for that. First, there isn't usually pressure to constantly attract new users, so there isn't as much pressure to change the UI to make it easier or simpler for new users, or follow UI trends. Secondly, the projects often don't have dedicated UI/UX designers who want to try out new ideas or make their mark, etc.
However, these aren't strictly good things, you can end up with an unintuitive UI with a steep learning curve that is hard to learn.
A computer is also a tool that does many things and the OS is arguably the important part of how a user wields that tool.
It does variations of the same thing.
A computer is more of a toolbox than a tool.
Perhaps this is why users dislike monthly SaaS - they cannot trust stability of the tool, because often the incentives are to keep adding features well past peak utility [ resulting in enshitification ]
On the other hand, if you hire a cook, then all you need to know is English (or whatever language they speak).
Even if the cooks keep changing, you can always just tell them "use the knife"
This is a huge complaint for me. I had to disable JetBrains from automatic updates because they wont stop trying to force their horrendous new UI on me, and the "Classic UI" plugin I have to use to keep my IDE working the way I have been used to for the past 15 years is never updated at the same time.
1: https://www.joelonsoftware.com/2000/04/10/controlling-your-e...
I just patched Claude Code's system prompts, pinned the version and stopped upgrading without first dissecting and auditing the executable. Even discovered Anthropic can remotely inject strings into the system prompt via some "growth book" or something. Neutralized that too.
Things got a lot better after I started doing this. It straight up fixed Opus 4.6, and Opus 4.8 got more consistent in my subjective experience.
Sadly there's nothing I can do about Anthropic's server side "system reminders" whenever some prompt trips their classifiers or whatever.
I'm in the process of switching to OpenAI and Codex. The open source harness is a breath of fresh air. We'll see how that goes.
[0]: https://pi.dev
Perhaps this[0] Pi documentation could help.
0 - https://pi.dev/docs/latest/models#anthropic-messages-compati...
I still haven't gotten around to trying it, though, so I don't know what the reality is. And new open-weight models have been released since then that I'd want to evaluate...
If I have too much usage I'll simply find something for the AI to do. Guided router update. Router hardening. Laptop hardening. Penetration test my router now that I somehow got into the cyber program. Local models research. PC build research. Smartphone research. Financial investments research. Business research. Medical research. Laptop firmware reverse engineering. Old video game reverse engineering. Pick a random open source project and let's explore the code base. Analyze all of my ten thousand HN posts and tell me interesting facts.
If I can't think of anything I default to picking a random git repository and launching a massive parallel code review session. That's guaranteed to kill any remaining usage in exchange for useful output. Then I can go enjoy my weekend guilt free. Unless they reset the usage.
GPT is also better than Claude from my personal experience. GLM 5.2 is equal and Kimi K3 is better.
But I mostly use Deepseek and pay api prices.
There is some movement in this direction--there is a new 'gh' subcommand called repo read-file for example, that lets agents view a file without cloning a repo. And I made something called venetianblinds that shows equidistant samples of a file. In combo they work pretty well:
gh repo read-file sqlite3.c --repo clibs/sqlite --output sqlite3.c && npx github:firasd/venetianblinds sqlite3.c
Shell execution and raw DOM access are great for a backstop, but you can go so much further with just a little bit of translation and delegation around the environment.
I think browser automation is probably the most apt scenario. Often a human who understands how a page is meant to be perceived can transform a megabyte of raw web content down to a few hundred bytes of plaintext without any reduction in fidelity. This can be achieved using deterministic code that is guaranteed to provide a perfect transform every time.
The performance difference between raw DOM access and curated plaintext is like a step function. With raw access you get maybe 10-15 steps into a complex workflow before the wheels pop off. With curated access I've seen it go 100+ screens without issues.
Simply managing the token bloat is probably the most important objective here. If that's all you focus on it will probably go really well.
Bash is a great control surface anyway for LLMs as it is wordy and powerful.
Another venetianblinds survey, of the Paul Graham 'What I Worked On' article that's used in many LlamaIndex examples:
npx github:firasd/venetianblinds pgworkedon.txt
Having a scripting language as a tool is powerfull, and can help remove unnecessary stuff from the LLM's context.
Incredible, in a way, that you still see it as a brand. I honestly thought it was a 10 years old blog post somehow making it to frontpage (as some old blog posts sometimes do on HN).
Who still uses DeadOverflow-full-of-outdated-answers? (outdated and often just plain wrong too)
I haven't been there in like 15 years or something. Feels like ages.
It's owned by some really stupid private equity firm since about 2019, which is trying to revive it as a platform for agents to talk to each other.
This is ridiculously complex to just type a few paragraphs. Nobody in their right mind would invest this amount of yak shaving... and yet I do so because I bet, rather confidently, that in few years, heck few decades, all those tools will be different (or maybe not, I still use Vim on my server, desktop but even mobile phone) but the lessons will remain practical.
IMHO the trust comes from trust yes but also more directly plain ownership.
Thanks for the heads-up that such a thing is possible. I will definitely be investigating it.
To summarize:
- start making a list of the manual tasks you do
- if those tasks involve running command line tools, add an item with the commands you run
- if they are manual tasks, add those too
- over time, keep automating one portion of the list at a time e.g. the commands can become a script, the manual tasks can become tickets to another team to automate etc
At the end of the above process you have a series of automated steps that become a system instead of a bunch of items in someone's head.
In my mind, the only thing that changed with LLMs is that it's faster to create the scripts and some of the manual tasks can be done by the LLM until you get a script to do that too.
We invented code to help do "mechanical" tasks over and over again in the same way. Why replace that with agentic systems??
P.S. This is also why the whole "just commit the prompt, bro" is such utter hogwash
0 - https://queue.acm.org/detail.cfm?id=3197520
The first case of large scale automation demonstrates this very well. Medieval manuscripts looked like this [0]. You can see the imperfections, but it's still a beautiful book. Gutenberg bibles omitted the scribe in favor of automated printing, but humans remained involved with the creative details in illumination and rubrication. The result is genuinely beautiful [1]. The works that were printed a century later rarely featured this kind of post-print human involvement [2]. This is still better than most of what's printed today, but it's a clear step down.
Now, there's a reasonable argument to be made that the quality differences don't matter for books and certainly don't outweigh the cost advantages. But imagine someone at your passport office has taken A.B.A. to heart and automated 90% of the job. Over time, the organization will stop handling all the edge cases and. But the edge cases didn't stop happening, they're just no longer visible. Beyond a certain scale (e.g. Google and other FAANGs), those inevitable failures manifest as seemingly capricious behavior that makes everyone hate your system and produces outcomes no human involved actually wants.
[0] https://i0.wp.com/blogs.princeton.edu/notabilia/wp-content/u...
[1] https://ff-65a4.kxcdn.com/assets/uploads/OriginalDocs_old/98...
[2] https://www.prepressure.com/images/Nieuwe-Tijdinghe-newspape...
For example, I use Codex and Claude Code by default, but when I need to look at the API surface, read tests, etc I have those tools setup to open Zed. Zed is also rapidly evolving in the other direction, where it's closer to the tools that are opening it. It won't be long, I think, until I can continue my prompt from inside Zed.
Trust lost is rarely ever regained. That they publish a piece like this on the Stack Overflow blog is insult to injury.
Now agentic coding produces the full 100 line change, and SO says the real crisis is that nobody has reviewed it properly. SO was fully responsible for creating the culture where no one wanted to use it.
It's not done when you can't trust it.
And shelling out to a cli? That's a pretty damn compelling and composable boundary, IMO.
This makes it unreliable and preferences are things I need to assume are treated as exactly that, preferences, not hard settings.
It is an area where it is more like working with an unreliable human than I would prefer.
> "Developers become attached to tools like Vim, Emacs, or an IDE because years of experience make those tools predictable extensions of their thinking. The attachment is less about features and more about accumulated trust, muscle memory, and a workflow built around known boundaries.
> AI coding agents disrupt that trust because they are fast but probabilistic, opaque, constantly changing, and capable of producing more code than humans can realistically review. This shifts the bottleneck from writing code to specifying, reviewing, validating, and operating it safely."
(Note the > is paraphrasing)
Trust is a big problem I'm having with these tools so far. What I've been running into a lot is, I'll get the equivalent 40 hours of work done in 8 hours, and I'm like, wow, that really was quick. Then I'll start using the application I'm making more directly (a tool for writing), and I'll start to see that it's broken all over the place in very surprising ways (ie, updating this menu item broke something on the other side of the app, etc.). So then I spend another 40 hours of real wall clock time kind of fixing everything that was broken, and at the end of those two weeks I'm like, did I actually go much faster or was that all kind of a wash? Because if I'm not going faster in overall terms, then the loss of deep understanding of the code base might not be worth it if my pace is the same.
I'm sure someone is going to be like "BRUH AUTOMATED TESTS" or "BRUH MODEL CHOICE". I have a LOT of automated tests, and I don't like fussing with models so I pretty much use Opus on high reasoning for most things (or the equivalent from other providers). Code review also doesn't help that much, for much of the same reason it doesn't tend to help find bugs in human written code either.. you're reading the happy path usually.
Anyway I wouldn't say these tools aren't useful, but, I'm deeply skeptical of all the productivity claims because I think people just look at one dimension of it while ignoring all the other important dimensions. Yeah you can generate a lot of crap fast, but most of it is not shippable and making it shippable does take time.
Also, is there ways for AI to help prevent the loss of deep understanding of your code base without you having to know every line of code deeply?
Am I giving the agent enough context? Well, I'm giving it as much as I can. Each submodule has an AGENTS.md, I have the agents add gotchas and instructions for some feature work when I discover where an agent went wrong, the codebase has a lot of comments along the lines of "if you edit this section, you need to also edit XYZ", and I lean on the type system as much as I can to make wrong-code not compile. It has access to playwright for driving the UI if it wants to. (Weirdly, I've found that Claude is really inconsistent about using these tools -- even though the instructions make it clear that it's allowed and encouraged. I think if your workflow differs from the models training and thusly you have to tell it so in AGENTS.md/CLAUDE.md, then it's very inconsistent about following those instructions. For instance, I don't want Claude to commit and I don't want it to sign commit messages, and it still does that all the time even though it's my like #1 directive of "don't mess with my git history")
There are some things though that are very hard for it to test. I'm exporting essentially a programming language to three game engine runtimes. They all have automated tests, but, I think people that have worked in video games know that games are very hard to automate testing on. This isn't really the fault of the agent I would say, just the nature of the problem, but it is worth noting.
I guess this is a long winded way of saying, even with LLMs tech debt is a thing you have to manage, and I think managing tech debt becomes even more important when you're dealing with LLMs, not less important.
He says "40 hours" because he feels like he's managed to do 40 hours worth of work in this time, but then has to spend another "40 hours" (actually: 1 day) just going around kicking tyres.
Obviously the implication is that it's a net gain of some kind, but he's unsure if he caught everything.
(sorry to reiterate the GP, but I feel like you missed the important nuance that it's not a real 40 hours of time).
Alice: "in the end, AI doesn't make me any faster because it still takes 80 hours to do 80 hours of work once I fix it"
Bob (AI booster): "actually you might've been holding it wrong, did you try XYZ?"
Carol (super AI booster): "Bob, actually Alice means it took 16 hours to do 80 hours of work. So it did work for her."
Alice: "no I fucking didn't"
I would say that this means that it only appeared that you got 40 hours of work done.
> I'm sure someone is going to be like "BRUH AUTOMATED TESTS" or "BRUH MODEL CHOICE". I have a LOT of automated tests, and I don't like fussing with models so I pretty much use Opus on high reasoning for most things (or the equivalent from other providers).
Bruh, manual tests. It shouldn't be an entire day until you "start using the application more directly". It's rare that you really know what you want until you try something that isn't what you want, and figure out what's wrong with it. LLMs aren't privy to your thought patterns, so they can't help with that.
And I can hear your objection now: correct, it's probably not worth all that for a throwaway, but the effort per output goes down as the infra builds up and you end up with a program that can reliably expand.
I'll also just point out my philosophy for using LLMs for this project, which is that I'm not trying to go as fast as I can. (I want to go at a good pace, but this isn't an experiment to just finish something over a weekend). The 150k LOC have come about since February, with some mix of me writing code and LLMs, so on average I'm probably bringing in about 800 LOC per day, which I imagine a lot of vibers would find to be glacial. To me that's the sustainable rate of what I can do when you factor in that I need to test drive every feature, make sure it doesn't conflict with another feature, check for bugs, check that the code looks reasonable, and debugging. (I also think that rate limit is specific to this project: I could see easier to test things going much faster, and harder to test things going slower)
I further suppose you are not writing the test cases in its whole by hand. But you try to specify them before you hand out the development task to the LLM, right?
It sounds like you just introduces a new team member which is not trustworthy yet. Normally you would review each change of him and explain how to improve hisself and the code. But that's not possible to a fixed-state LLM. That sounds exhausting.
If your Markdown Files should aim at improving the LLM contributions, you are again stuck with the fact that it did not follow in the first place.
So I conclude: You pay for an Intern who is not trustworthy and pay additional input token on Markdown Files to still no be certain about future contributions.
I just don't grasp how we as engineers are accepting this and integrate it into our craft. And: Everybody using External LLM Services, possibly providing the entire project as context, is allowed to let the source code of your company be leaked to some third party. I hope that party is trustworthy and does not have a track record of copyright infridgement. Because this would be fairly naive and reason to be fired. So we as Engineers knowing the implications should therefore point to these issues at the correct management level.
At least that's what I am doing shrugs
My own stance is that there's never any reason to go fast on anything. Communication has always been the bottleneck. Whether it's about gathering requirements or understanding the purpose of a badly written code, any speed improvements I get has always been a small percentage of the overall progress.
What has helped more is my understanding of the platform and some theoretical knowledge. Because one I get the information, I can quickly derive a solution in my mind. And that solution has always been easy and fast to implement, at least the happy path. 90% of the time taken in coding is always about handling all the edge cases, aka fixing bugs. And writing tests so that you're not easily introducing more bugs.
I update my toolset all the time. It always results in discomfort and backlash but people don't understand that the goal isn't their perception or trust. It's about skill, ability, and execution. This idea probably won't get me promoted but it will get me paid. I am not attached to tools because I learned the hard way that they will always find a way to take them from me.
Zero trust is a better alternative for people like me. In terms of cybersecurity, being attached to a tool is crutch because fatal flaws in every design are frequently found. As it relates to agentic AI, I never select the "Yes, trust the AI and let Claude execute arbitrary commands in a non-sandboxed environment" option. However, I frequently utilize agents, but I'm not going to have the "Jesus, take the wheel" moment with them right now. That being said, AI is a very helpful tool that helps me create boilerplate code, brainstorm ideas, and review my work. I also anticipate when AI can, in fact, take the wheel and I'm looking forward to it.
Parallel to this, I also know that developers often disagree, and I'm not casting judgement on anyone for being attached. If it's Turing-complete, then I have the background to complete the task. In these scenarios, I just adopt whatever tools work best in team building because, in my own words, I'm not too attached to the way I do things.
Don't we all?
(I can see that you are starting to get downvoted as well. Spread the love.)
Stack overflow was interesting. Its design was the only thing like it at the time, and made it a Schelling point for programming knowledge distribution, but also a welcoming environment for the programming equivalent of grammar nazis.
Some of those programming nazis, of course, had suffered at the hands of previous ones on stack overflow before becoming "enlightened." And thus, the generational hazing began.
It was great if google directed you to exactly the right answer, but god help you if you couldn't figure it out, and posed a question that someone thought didn't contain an MCVE.
Also, a few too many of the high-reputation people would post complete garbage on topics they knew absolutely nothing about.
And "grammar nazi" has been around since at least 1990, and the infamous Seinfeld Soup Nazi since 1995.
I'm not sure of the etymology of the "not everybody's a Nazi Nazi" but I'm sure you're not the first.
But in any case:
> You can't just call everyone you don't like a Nazi.
Yes, yes, I can. I don't (because I reserve the appellation for certain particular kinds of attitudes), but I could if I wanted to.
Here is the proof