Hacker Newsnew | past | comments | ask | show | jobs | submit | xianshou's commentslogin

the future is here and one should be thankful for its slightly uneven distribution. otherwise we would hardly have anything left about which to develop strong opinions!


Really missed the opportunity for "Airemin".


A lovely example of a study that is both obviously true and misses the point.

Music with lyrics directly interferes with any task that has a verbal component, and the worse you are at multitasking, the worse the interference. Despite being terrible at multitasking, I still listen to music with lyrics. Why? Principally because the alternative, hearing all the conversations in my immediate vicinity, is usually both more distracting and less pleasant. But there are also auxiliary benefits, such as an increase in "work stamina" and a passive signal to coworkers to interrupt only if it's important.

Now, I could listen to lo-fi all day, or three-hour soundtracks on Youtube, and sometimes do, but it gets boring pretty fast!

Anyway: obviously true, still worth it because the alternative is worse.

(By the way, other mitigating strategies: listening to music in a language you don't understand, or listening to lyrics so familiar you can screen them out. My top Spotify songs all get played several hundred times a year.)


Haha, listening to music in the other language I understand (from the language required for the task) helps in my case. ;)


Even as someone extremely firmly on the other side of the AI debate, I must appreciate the craft.

Now, to give Claude the steganogravy skill...


Maybe call it steakanography so it stands out from mere steganogravy.


From the file: "Answer is always line 1. Reasoning comes after, never before."

LLMs are autoregressive (filling in the completion of what came before), so you'd better have thinking mode on or the "reasoning" is pure confirmation bias seeded by the answer that gets locked in via the first output tokens.


Yeah this seems to be a very bad idea. Seems like the author had the right idea, but the wrong way of implementing it.

There are a few papers actually that describe how to get faster results and more economic sessions by instructing the LLM how to compress its thinking (“CCoT” is a paper that I remember, compressed chain of thought). It basically tells the model to think like “a -> b”. There’s loss in quality, though, but not too much.

https://arxiv.org/abs/2412.13171


For the more important sessions, I like to have it revise the plan with a generic prompt (e.g. "perform a sanity check") just so that it can take another pass on the beginning portion of the plan with the benefit of additional context that it had reasoned out by the end of the first draft.


Is this true? Non-reasoning LLMs are autoregressive. Reasoning LLMs can emit thousands of reasoning tokens before "line 1" where they write the answer.


They are all autoregressive. They have just been trained to emit thinking tokens like any other tokens.


reasoning is just more tokens that come out first wrapped in <thinking></thinking>


there are no reasoning LLMs.


This is an interesting denial of reality.


A "reasoning" LLM is just an LLM that's been instructed or trained to start every response with some text wrapped in <BEGIN_REASONING></END_REASONING> or similar. The UI may show or obscure this part. Then when the model decides to give its "real" response, it has all that reasoning text in its context window, helping it generate a better answer.


I don't think Claude Code offers no thinking as an option. I'm seeing "low" thinking as the minimum.


Ugh. Dictated with such confidence. My god, I hate this LLMism the most. "Some directive. Always this, never that."


I appreciate not having to read this guy again.


Great work! Why no benchmarks though?


Nice! 5 bucks says you can swap this in for your average software kanban and it does a better job.


Safer than clawdbot/moltbot, I'll bet.


What makes you think it isn’t clawdbot under the hood?


it's not :)


Why? It seems just as likely to follow prompt injection commands.


Incidentally, Chroma also produced the single best study on long-context degradation that I've come across:

https://research.trychroma.com/context-rot

Before that, I cited nolima (https://www.reddit.com/r/LocalLLaMA/comments/1io3hn2/nolima_...) constantly to illustrate how difficult tasks involving reasoning or multi-step information gathering degraded much faster than the needle-in-haystack benchmarks cited by the major labs. Now Chroma is the first stop. Nice job on the research!


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: