5. been using gemini,claude code。底层模型的能力的发展实在令人惊讶 (and ppl telling you that models are pretty much the same and end users can't tell the difference are not being honest or don't know what they are talking about). because models are so good now, it needs less 'harness' (programming logic), which does as much to limit its abilities as to prevent it from doing the wrong things. with every model update, the Claude Code team removes code rather than adding to it.
这意味着什么?things that have clear right/wrong answers are certain to be automated by LLMs, from coding to solving math puzzles to ... stock picking? why not? the LLM prediction is going to be either right or wrong. for now you as an analyst might fancy elaborate prompts and handcrafted workflows and thinking it might give you an edge. perhaps for now. but the model will become so good at it that your clever prompts are going to hold it back, like human 棋谱 were holding back AlphaZero.
Maybe what's left for humans to do, and the only place where tasteful prompts still have a place, is where there are no clear right/wrong answers. but even there i am not so sure... maybe we can't directly measure if a story is any good, but we can measure proxies, e.g. how viral did it go? when you measure by proxy there is always the chance of "reward hacking"...
what's the takeaway? (self prompting here) -- keep in mind that models will advance to a point where no 'harness' is needed, at least for easily verifiable tasks. For the remaining tasks, you can almost certainly achieve the desired results by good prompting -- not to handhold models with detailed workflow, but get across your taste, aethetics, vision.