
LLMs Reward Expertise: Why Deep Domain Knowledge Matters More as AI Gets Smarter
- luandnh
- Ai tools , System design
- October 3, 2026
Table of Contents
The old CSS gap and the new backend gap
Back in the 2010s, if you had a skill gap - say, you couldn’t write CSS - you had exactly two options: lean on a more skilled colleague, or pray someone had posted the answer to your precise problem on Stack Overflow.
Not anymore. Today anyone can write sort-of-okay CSS by delegating the task to an LLM. LLMs turn everybody into a generalist.
Which is why a lot of people don’t think there’s any skill involved in working with LLMs. Want the product LLMs can deliver - PhD-level math, code that runs but sometimes has no taste, LinkedIn-style prose that makes you cringe? Just ask. Since everyone is talking to the same models, “skilled prompters” get the same results as people touching an LLM for the first time.
That’s wrong. The most important skill in prompting is expertise in the domain you’re prompting for.
Sean Goedecke, an engineer at GitHub, wrote “LLMs reward expertise” in July 2026. It hit the top of Hacker News at 1416 points with nearly 600 comments. Not because it promised anything new, but because it named something many people felt but couldn’t articulate: the stronger AI gets, the wider the gap between the strong and the weak in any given domain, not narrower.
Terence Tao and the ChatGPT I’ll never touch
The best illustration is Terence Tao’s conversation with ChatGPT about the recently discovered counterexample to the Jacobian Conjecture.
Quick context for anyone not following. The Jacobian Conjecture says, roughly: if a polynomial map over the complex numbers has a Jacobian that is a nonzero constant (i.e. it’s locally invertible everywhere), then it’s globally invertible, with a polynomial inverse. People believed it for over 80 years. In July 2026, someone found a counterexample in three dimensions - using an AI system called Fable - and the conjecture collapsed in dimension 3 and up (dimension 2 is still open).
Tao wrote a “digestion” post to work through that counterexample with basically no algebraic geometry. Part of that process was a dialogue with ChatGPT.
Sean watched the conversation and pulled out a few traits of how Tao prompts:
- His messages are very short and to the point. He doesn’t respond point-by-point to the model, just to the gist.
- The model’s outputs are far more concise than when an ordinary person like Sean asks GPT-5.6 about math. By signalling expertise, Tao shunts the model into “talking-to-mathematicians” mode instead of “explaining-to-amateurs” mode.
- He pushes back when output looks wrong, but never directly contradicts. He says things like “this looks more complex than I was hoping for,” not “that’s wrong.”
- He makes the leaps himself. He almost never takes the model’s advice about where to go next.
But here’s the crux, and the part every “10 expert prompting tricks” article skips: you cannot prompt like Tao just by copying those four tricks. The key is actually understanding the math - pulling the relevant idea out of ChatGPT’s multi-paragraph response, suggesting alternate approaches or formulations, and identifying what looks weird.
Sean puts it bluntly: “Terence Tao is a better mathematician than I am a programmer.” But the mechanism is identical in his own work. If you have a good theory of your codebase - a solid mental model of the system you own - you can push the LLM much harder than if you’re flying blind. Because you have your own sense of what a good solution might look like, you can say: “no, I think this could be simpler here,” or “but don’t we already do X?”, or “can we express this problem in familiar terms?”
Two loops: one spins, one converges
To see the mechanism clearly, picture two prompting loops. This is where “expertise” stops being a vague feeling and becomes a concrete feedback loop.
graph LR
subgraph "Non-Expert Loop"
N1["`Vague prompt
'it's throwing some error'`"] --> N2["`Model answers at length
no way to judge right/wrong`"]
N2 --> N3["`Read, try, get lost
ask for another feature`"]
N3 --> N1
end
subgraph "Expert Loop"
E1["`Short prompt, correct terms
'data race on this map, use errgroup'`"] --> E2["`Concise answer
right expert register`"]
E2 --> E3["`Flag what 'looks weird'
propose your own direction`"]
E3 --> E1
end
The left loop never converges. You have no criterion for whether the output is right or wrong, so each iteration just adds confusion. The right loop converges, because each pass you inject a constraint grounded in real understanding.
There’s a Hacker News comment that nails the left loop. A user named krisoft tells how a friend of his (no software engineering background) wanted to build a simple single-page web app. He told her to try building it with an LLM while he watched. He expected the AI would have no trouble writing the code, and was just curious whether it would realize she was a novice.
The result: she didn’t have the vocabulary to ask the AI to write code. She went around in circles brainstorming features, each idea getting more complicated. After an hour and a half and a huge number of messages, they killed the experiment. Meanwhile krisoft, who knows the terminology, would have needed one message: “Write me an HTML page that does X, Y, Z.” The only difference was vocabulary.
That’s the painful insight: a gap in technical vocabulary blocks the feedback loop from closing. You don’t know how to ask “how do I build what I want,” so you ask “what should I do next,” and you fall into the vortex.
Backend translation: technical vocabulary is money
Before LLMs, a backend engineer who knew the terminology cold - “race condition,” “critical section,” “compare-and-swap,” “backpressure,” “head-of-line blocking” - had an edge. With LLMs, that edge doesn’t disappear. It multiplies.
Take a concrete example everyone hits: fetching from N URLs concurrently.
Someone without the context, goroutine, and sync vocabulary prompts: “Write Go code to fetch many URLs at once, fast.” The LLM returns something like:
func FetchAll(urls []string) []string {
var results []string
for _, u := range urls {
go func(u string) {
resp, _ := http.Get(u) // no timeout, no context
body, _ := io.ReadAll(resp.Body)
results = append(results, string(body)) // data race
}(u)
}
return results // returns immediately; goroutine leak
}
This code looks right. It runs. It returns data. And it’s wrong in at least five ways: a data race on results, unbounded goroutines (N = 10,000 URLs and you’ve shot yourself in the foot), no timeout, no status-code handling, and a full body leak because nothing is closed.
An engineer with the expertise spots all five instantly. More importantly: he knows how to turn that expertise into constraints in the prompt. Not “write me concurrent fetch code,” but “use errgroup.WithContext, cap concurrency at X, honor context cancellation, don’t leak goroutines.”
- Go (Naive)
- Go (Expert Guided)
// Prompt: "Write Go code to fetch many URLs concurrently"
// What you get when the user states no constraints.
func FetchAll(urls []string) []string {
var results []string
for _, u := range urls {
go func(u string) {
resp, _ := http.Get(u)
body, _ := io.ReadAll(resp.Body)
results = append(results, string(body)) // DATA RACE
}(u)
}
return results // returns before goroutines finish; leaks
}
// Prompt: "errgroup.WithContext, bounded concurrency, honor ctx
// cancellation, pre-allocate, cap body size, no leaked goroutines"
func FetchAll(ctx context.Context, client *http.Client, urls []string, limit int) ([]string, error) {
g, ctx := errgroup.WithContext(ctx)
g.SetLimit(limit) // bounded concurrency, no goroutine storm
results := make([]string, len(urls)) // one allocation, no append growth
for i, u := range urls {
i, u := i, u
g.Go(func() error {
req, err := http.NewRequestWithContext(ctx, http.MethodGet, u, nil)
if err != nil {
return fmt.Errorf("build request %s: %w", u, err)
}
resp, err := client.Do(req) // client carries a Timeout
if err != nil {
return fmt.Errorf("get %s: %w", u, err)
}
defer resp.Body.Close()
if resp.StatusCode != http.StatusOK {
return fmt.Errorf("get %s: status %d", u, resp.StatusCode)
}
body, err := io.ReadAll(io.LimitReader(resp.Body, maxBody))
if err != nil {
return fmt.Errorf("read %s: %w", u, err)
}
results[i] = string(body) // distinct index -> no race
return nil
})
}
if err := g.Wait(); err != nil {
return nil, err
}
return results, nil
}
Both snippets are “AI-written.” But the second only exists because the person prompting knew what they needed. The LLM didn’t invent g.SetLimit, didn’t add io.LimitReader, didn’t remember http.NewRequestWithContext. It knows all of them. It just didn’t know you needed them.
Tip
Practical rule: when prompting Go code, name the specific libraries and primitives (errgroup, context.WithTimeout, sync.Pool, semaphore.Weighted) instead of describing intent in the abstract. You’re not teaching the AI to code - you’re narrowing its search space to the region you want.
pprof: where expertise pays off hardest
There’s a class of question non-experts never know to ask: questions about evidence.
A junior facing a slow endpoint prompts: “How do I make Go code faster?” The LLM returns a generic list: use caching, avoid allocation, use goroutines, cut syscalls. Sounds reasonable. It’s all guesswork.
A senior prompts differently: “Here’s my pprof profile. renderMetric is 43% of CPU. Here’s the source. Optimize it without changing behavior.” He isn’t asking “how do I go faster” - he’s asking about a specific function, already measured with real data.
import (
"log"
"net/http"
_ "net/http/pprof" // registers handlers on DefaultServeMux
)
// Expose the debug endpoint on its own port. Never expose it publicly.
func startDebugServer() {
go func() {
log.Println(http.ListenAndServe("localhost:6060", nil))
}()
}
Then pull the profile:
go tool pprof -top -cum http://localhost:6060/debug/pprof/profile?seconds=30
go tool pprof -http=:8081 cpu.prof # inspect the flame graph
Note
The difference isn’t that the senior knows go tool pprof. The LLM knows it too. The difference is that the senior knows you don’t optimize without measuring - and knows to hand the LLM a concrete profile with a “don’t change behavior” constraint, instead of letting it guess.
The same goes for memory allocation. People with the expertise understand that append without pre-allocation copies repeatedly, that interface boxing creates hidden allocations, that sync.Pool makes sense for short-lived high-frequency objects but carries its own cost. They turn that understanding into a prompt:
// Naive: append without pre-allocation -> repeated realloc + copy
func Process(items []Item) []Result {
var out []Result
for _, it := range items {
out = append(out, transform(it)) // reallocates each time cap is exceeded
}
return out
}
// Expert guided: allocate exactly once, pass a pointer to avoid big struct copies
func Process(items []Item) []Result {
out := make([]Result, 0, len(items)) // one allocation, exact capacity
for i := range items {
out = append(out, transform(&items[i])) // avoids copying large Item
}
return out
}
These two versions can differ by 30-50% runtime on a hot path. Someone who doesn’t know won’t think of it; the LLM won’t offer it because it doesn’t know your data sizes.
The bigger picture: the bottleneck moved to the human
Put it all together and Sean lands on a conclusion I think holds for backend and every other domain:
For many tasks, the human is the bottleneck, not the model, because the difficult part is communicating to the model exactly what kind of solution the human wants. The information is “in the model” already, but it takes a very smart human to pull it out.
Picture it as a funnel.
flowchart TD
subgraph "Before LLMs"
A["`Syntax & APIs
Needed someone who could write it`"] --> B["`System design
Needed someone who knew what they wanted`"]
end
subgraph "After LLMs"
C["`Syntax & APIs
LLM does most of it`"] --> D["`Prompt + judge output
NEW bottleneck: expressing intent
and spotting what's wrong`"]
end
Historically, most of an engineer’s value sat at the bottom of the funnel: knowing syntax, knowing APIs, producing code that runs. LLMs ate nearly all of that layer. But once that layer is automated, value shifts up: knowing what you want, judging output, sensing what “looks weird.” And that upper layer is precisely the part that cannot be automated by copying prompting tips, because it requires real system understanding.
This also matches something I’ve written about before: system design problems are dominated by concrete specifics, not generic principles. Knowing CAP Theorem or the Saga Pattern is fine, but reading your own codebase - which service calls which, which transaction holds which lock - is what lets you ask the LLM the kind of specific question Tao asked ChatGPT about his own math.
The flip side: “reassuring” arguments and the anecdote trap
If the post stopped here, it would sound too comfortable. And that’s exactly what I need to be suspicious of.
On Hacker News, several commenters flagged the intellectual hazard in Sean’s thesis: it reassures a specific group - developers - and we tend to believe arguments that make us feel safe about our future. I agree. “Expertise still matters” is an easy belief to accept because it’s pleasant. But its odds of being true aren’t higher just because it’s pleasant.
Second, the evidence for “expertise wins” is largely anecdotal, and anecdotes cut both ways. krisoft’s comment (the friend who couldn’t prompt) sits next to a comment from nvrmndmnm with the opposite result: his girlfriend, a hair stylist with zero coding background, installed Arch Linux + Hyprland, riced the whole setup, got Steam and Portal running, and then built a working Telegram bot on free-tier Gemini plus a Kimi key. Zero expertise, real result.
So what separates the two cases? Commenter cgannett nails it: krisoft’s friend lacked the vocabulary to ask, but juicing that vocabulary out of the AI is itself a skill. nvrmndmnm’s girlfriend isn’t an expert, but she knew to ask “what are web pages built from” and “can we use those building blocks to build our own app.” That’s a tinkerer’s disposition - curiosity plus the habit of asking the right questions - something engineers treat as obvious but which is really a culture, a practice, not an instinct.
And then the strongest counterpoint: OpenAI (and other labs) found those math results with non-expert prompts - they didn’t need expertise for the model to propose discoveries. Sean’s reply lands: OpenAI has a team of expert mathematicians that checked and filtered the model’s suggestions, and you can’t currently skip that step. Put another way, expertise doesn’t have to sit in the prompter - but it has to sit somewhere in the pipeline for the number to be believable. That’s the subtle point: the question isn’t “does AI need experts,” it’s “which stage needs the expert.”
So what do you do with this
If the thesis holds, a few consequences follow:
Stop studying prompt engineering; go deeper on your domain. “10 advanced prompting techniques” courses are low-value. What pays is understanding the system you own deeply enough to ask questions the AI doesn’t know how to answer.
Build a theory of the codebase. Not memorizing code - holding a solid mental model of data flow, implicit constraints, invariants. With that, you can use Tao’s move: propose your own direction and refuse the model’s advice.
Use AI as a sparring partner, not an oracle. Tao doesn’t take ChatGPT’s advice on where to go next. He uses it as a wall to throw ideas against. In backend terms: you propose a design, the AI critiques it, you judge the critique, you decide. The judging role is never delegated.
Where you’re not the expert, admit it and route around. Sean is explicit: everyone mixes both modes, because we have expertise in some areas and are blind in others. When you’re blind, you cling to the LLM and try to notice when you’re in the vortex - and when you do, the right move is to find a human expert, not to ask the model forty more times.
Finally, keep some skepticism about the attitude. By the time we have enough research to assert “expertise still wins,” the landscape may have shifted under our feet again. But until then, the evidence - anecdotal as it is - leans one way: with the same model, the person who understands the system asks better questions, and better questions produce better results. AI doesn’t flatten the playing field. It just moves where the weight sits.
Warning
Don’t use this post to reassure yourself (“I’m a senior, I’m safe”). Reassuring arguments are the easy ones to believe. Read it as a hypothesis to test on yourself: next time you catch yourself stuck in a brainstorm loop with an AI, ask whether you’re missing expertise - or just missing a human expert.


