I am not familiar with the standards of publishing in machine learning, but as someone trained in a mathematics background, this paper seems relatively light on details and heavy on exposition. Is that typical? Is this a really novel idea? Not trying to be snarky, just trying to understand how meaningful this is.
It's cool that it proves that a bunch of vectorized outputs from an unknown embedder on an unknown dataset is in no way private, because of this ability to reverse engineer the embedder.
I talked to the author at his poster session at neurips and was able to get the gist, though I had read a lot about the platonic representation hypothesis, and this was one of my top 10 favorite papers in the conference.
They link to their code on GitHub - see footnote 2 on page 2. I don't see it linked anywhere else, which makes it easy to miss. https://github.com/rjha18/vec2vec/
You are not wrong. But this has by no means proven its up to the standard of being publishable in a machine learning journal. Its on arXiv.org, which, lets face it, at the end of the day is a vanity press.
Calling the arxiv a vanity press shows you aren't a researcher. In math and physics all the best stuff is on the arxiv and the general level is well above the level of most journals. Journals mainly serve as accreditation and many are basically mediocre - the review process resta more value than it adds overall.
What is your definition of a vanity press? Mine is that it will publish anything from anybody.
Sure, there's lots of good stuff on there. But there's lots which is not good. Publishing there is not doing science. Science requires peer review.
Its a wonderful resource, but when overworked journalists (or worse, AI) just picks up a paper from there and writes a breathless article on how cool this new idea is, it is misleading to the general public.
Worse, it can get the general public (and from the comments here, even people working in the field) to forget what science is and why arXiv is not a scientific journal.
In math and theoretical physics the arxiv is the default place to publish and where one goes to find the best and most current work. Papers take a long time to review (years) particularly hard and good papers and publication in journals mainly serves the needs of bean counting dean's and quality control morons in the administration. Actual scientific communication is via the arxiv. Self selection mitigates most ills you describe someone who regularly posts crap on the arxiv will do themselves damage professionally and disappear.
The pace of things is moving along so rapidly right now, I’m not sure that waiting for peer reviews is always a wise move. Doubly so if there’s a paywall; why limit your article’s impact by placing it where practitioners’ agents might not be able to access it? The rapid progress right now is challenging for conventional academic processes.
If the value of the paper is difficult to independently verify, for example, if it depends on the credibility of the author, then the academic ritual can add something. If it’s a mathematical result, one that can be automatically verified, or a machine learning technique that anyone can try with Claude code reconstructing it for them, this sort of pre-print publishing model is advantageous.
Because...science? It's not science until it passes peer review.
I'm not advocating that everybody stops posting to arXiv, and I'm not saying you can't find good stuff there. I'm just saying, it's a vanity press, there is absolutely no guarantee of the paper's quality.
And being published by a famous professor from a prestigious university is also no guarantee. If we've learned anything from the non-reproducibility crisis, it is that a paper's origin story is no guarantee.
We all know that science is about experiments and evidence whereas peer review is a social process.
Confusing science with peer review is very unscientific. Galileo’s work did not pass the dominant socio-scientific review process of his day and we all know how that turned out. Peer review serves a purpose but peer review is not science per se — unless, perhaps, your field is sociology, and you study the social processes of science practitioners, so that peer review outcomes are the evidence that you use to study those social processes.
Peer reviewers don't generally reproduce the work in the publications they are asked to review, at least not during the review process itself. Neither do journal editors.
If you want a real vanity press, check vixra and its origin story.
ArXiv does have moderators and endorsers in each area, although they are usually light-touch, weeding out literally unreadable submissions, ones that so obviously ignore formatting guidelines that it beggars belief they could comply with a typical journal's rules, and ones that are clearly submitted to the wrong area.
Having a stable document early (a pre-print) tends to widen the scrutiny of papers that may ultimately be published to a journal whose editor's expertise lies in a very different area from the submission; this can and does lead to author corrections being made before journal publication.
There are plenty of peer-reviewed papers which are hot garbage that have found their way into prestigious high-impact journals like Nature and Science (see <https://retractionwatch.com/the-retraction-watch-leaderboard...> for examples, and note none of the top 10 are in areas covered by the arXiv).
> being published by a famous professor from a prestigious university is also no guarantee
Everyone in academia knows this, including the vast majority of famous professors from prestigious universities, because of online repositories like <https://retractionwatch.com/retractions-by-nobel-prize-winne...> (and much more sadly because of <https://en.wikipedia.org/wiki/Nobel_disease>, lower-profile versions of which academic paper-writers -- and dissertation writers -- tend to encounter as they chase the history of the problem before them or read late citations to works they are relying upon). This particular part of your set of claims is is not a real problem in academia or with the arXiv in particular.
Finally, what value does publication in a predatory journal bring? Do you believe that peer review and editing were actually even performed in the majority of MDPI's most predatory journals, for example? <https://www.predatoryjournals.org/news/list-of-all-mdpi-pred...> Some papers published in some of their pay-for-publication open access journals don't even get submitted to the arXiv; one might hope this is out of embarrassment by the authors, although the (low) bar set by the arXiv itself is certainly a factor.
The remedy for a reasonably argued but wrong academic paper isn't lack of publication, failed peer review, or editorial alteration, but rather reply papers. That's the academic dialogue.
Thanks for this reply, your points are well taken and it really got me to think about this stuff.
I suppose you could view arxiv as being a way to let more people do a peer review on the papers, i.e. as one component of a more open and transparent peer review process. As such, it is certainly a welcome alternative to what the snootier journal reviews do, which is to put off reading your paper for a year, and then spend about 50 seconds looking for a reason to say "no".
But I had a narrow point to make when I said that arxiv is a vanity press: just because something is on there doesn't guarantee that it is correct, or even useful. It is a vanity press in the sense that it will publish pretty much anything--modulo obscenity and porn laws, and apparently also some sanity checking by arxiv moderators--but so does a vanity press for print books.
To your point of predatory journals, power-broker reviewers and such, yes, I have done my fair share of suffering from them. It slows down good research and passes through bad research. The system is badly in need of reform, and arxiv gives a much needed alternative.
But really, the problem is (as one of my profs said) that science is totally run on volunteer, unpaid labor. Peer reviewers are very busy, get paid nothing, it is mostly a distraction. This leads them to look for other ways of being compensated--e.g. the can try to be a power broker, a feared authority, etc.
The real solution would be to actually pay peer reviewers and fund them so they can actually reproduce the results. Here's how it could work: when a researcher submits a grant proposal, they budget for enough money to fund the research and enough money to peer review and reproduce it. If the research isn't good enough to warrant an attempt to validate it, then it isn't worth doing in the first place.
Thanks. Nobody ever admits when they are wrong on the internet, so I'm trying to be the change I want to see.
W.R.T. the grand historical tradition, etc, note that the great majority of people who practiced that were, like Plato, independently wealthy and otherwise idle. Those who didn't need any money.
That's why we give Judges, teachers, and Professors tenure, a guaranteed job for life. People are more likely to be fair judges of other people's performance when they don't need the money.
So I guess if you want a non-conservative fix, it would be universal guaranteed minimum income. Something we already have, BTW, but you have to be over 62 to get it.
We've tried to open it up to more than just the 1%, but we didn't really go all the way. It is really unreasonable to expect somebody who isn't independently wealthy to do free work. It is a lot of work to properly review a paper, and reproducing the results is something you simply can't do at all if you are supposed to already be working 70 hours a week generating your own research.
Of course a researcher paying reviewers would be a disastrous idea. But if the grant-awarding body would hire 3rd parties to validate that their money was actually well spent, I think it would be a win/win.
Thanks for the references on where paying reviewers has already been tried. I'd like to see others reproduce their success :-)
No paper is guaranteed to be completely accurate and correct at the time of publication; the remedy for honest errors is follow-on papers by the original authors or others in the field putting their names to formalized counter-arguments.
This is how the academic dialogue works, and it dates back to early durable records in the ancient world -- about as far before Plato as Plato's Academy is before today's universities. Innovations like academic journals come from "merely" four centuries ago, and pre-publication evaluation by peers about three hundred years ago. Anonymized pre-publication review is more recent still, and publishing schedules (as well as the needs of the submitter to have the submitted work dealt with in good time) don't admit a dialogue over some subtle point a referee takes some issue with. Moreover, such a pre-publication dialogue is generally non-public, and often also filtered through an editor. It can get messy [1].
As I'm sure you've discovered in your own post-secondary education, academia is very conservative (in the sense of "retaining what mostly works, even in the face of mounting scaling problems", foremost among those being the ever-increasing volume of publishable works generated by researchers).
Direct compensation of referees by journals has a number of pitfalls ranging from the fiddly details of how to deal with income tax liability pay (and credits towards future publication fees, etc) incurs to whether it compromises the objectivity of pre-publication evaluation (which is frankly already not very good in highly specialized areas).
Payment of reviewers by paper submitters is imho an even worse idea -- how do you preserve the anonymity of referees (especially with respect to the submitters)? Do submitters or reviewers have exposure to their (probably different) tax authorities for these transactions? (I can think of at least four national tax agencies that would view the received pay as taxable income, and also potentially subject to value-added tax on professional services. One already sees this problem with conference honoraria.)
Another -- one that already exists in the real world -- might be for institutions and grant-funders to build in support for peer review of others' work by project members. To the extent that detailed pre-publication review benefits academia, and societies that fund academia, as a whole, maybe referees should be rewarded (or even obliged) to do paid reviewing duty. One might compare the requirement to do pro bono work imposed on many practising lawyers by their regulating authorities (especially interesting are jurisdictions in which a party losing a court contest to a party represented pro bono usually must pay into a fund that supports the overall scheme).
Personally I think it would be more useful to everyone to have a solid reply paper published in due course than a blocked initial publication. My ideal is that referees aim to help submitters not accidentally professionally embarrass themselves.
Then on your reproduction idée fixe, supporting the generation of papers which confirm reproduction is something that societies could certainly work on. Many incentives to write and submit papers calling out flaws in others' work already exist, and of course could be expanded. I don't think any of that should be done with masking of the participants. Additionally, reproduction of results is often less interesting than (sometimes much) later testing of published results in significantly different ways.
(Indeed, foundational papers from more than ninety years ago are often re-proven by modern-day experiments, both as a side effect of (or opportunity arising from) pursuing the primary research goal, and as a way of testing e.g. new metrology technology, new computational tools, mathematical advances, and so on. Also, some of those foundational papers simply could not be tested with great accuracy with tools existing at the time of publication. Their publication, however, tended to drive the development of such tools, and in modern days such foundational papers are found to be good only within certain limits that could not have been explored closer to the time of publication.)
The entire thing is so generic and full of wise sounding things with no deep concrete content. It screams "an AI wrote this and a human thought it looked ok and pushed it".
Not OP, but LLMs didn't exactly invent the kind of content that's not worthy of your time, they merely turned it into a commodity (and made everyone have to work much harder figuring which is which).
I'm not an MBA over here, but this math seems wrong. If they are spending $240 in increased costs, then they only have to make about $247 in additional revenue from that spend to preserve a 3% margin. That seems much more reasonable if it increases the probability that customers find the product they are looking for and have a good experience.
No, because costco has a 3% margin. A cart of stuff at costco costing 247 will yield 240 to various operating costs, and roughly $7 actually to costco.
If you have a lemonade stand you might sell a cup for $1, but overall after paying yourself and for the cups, lemons, etc you might only get 3 cents each cup.
I understand that. The AI software is meant to be a productivity enhancer for the employees using it. Other than the licenses, some training etc, there are no operating costs associated with it. Just by using the software, I don't suddenly have to pay more for salaries, retirement plans, etc, which are things that in aggregate produce the 3% margin. Maybe I have to pay more in logistics because I'm moving more product now, but I think the point stands.
The point is if you are paying for software but not increasing your revenue by your margin you are lowering it even if you are increasing profit in absolute terms.
What you are mentioning with salaries is not relevant.
I think we can agree to disagree here. I don't see how a company needs to have an $8000 increase in revenue to justify a $240 software purchase. You are assuming that the current operating cost for every dollar of revenue is also applied to every incremental dollar in revenue gained from software efficiency, and that is just not true.
I agree here. OP is taking a retail company's entire profit margin, which includes a lot of operating costs, and estimating that the AI subscription will have the same margin. The AI subscription is software though, it probably has the operating costs and profit margins of software.
I interviewed someone recently who worked at Meta a couple years ago. He was a software engineer, was paid a bunch of money to mostly up dashboards all day, and eventually quit because it was neither interesting nor challenging.
It is not determined by the derivative, it's the antiderivative, as someone else mentioned. The derivative is the rate of change of a function. The "area under a curve" of the graph of a function measures how much the function is "accumulating", which is intuitively a sum of rates of change (taken to an infinitesimal limit).
This is some high quality content. Love the visual animations to go along with the mathematical ideas. Did a great job helping to tie the algebra to geometric intuition, but I think the importance of commutators could have gotten a little bit more exposition.
The disk model of hyberolic geometry is made to map hyperbolic 2 space (which is infinite in area) into the finite interior of the disk. In order to capture this, the normal euclidean notion of distance is distorted by a function which allows "distances" to go to infinity as a curve approaches the boundary of the disk.
Binary search minimizes the number of expected moves until you find the target. If you are already ahead, this is a natural thing to want to do. The reason why this doesn't work when you're behind is that your opponent can also do that and probabilistically maintain their lead.
I know that it minimizes the expected number of moves. But, the goal is to maximize the probability that you win in fewer moves than your opponent, not minimize the expected number of moves. Given that your opponent is playing some riskier strategy, it's not intuitively obvious to me that your optimal moves for those two objectives are the same.
If it helps your intuition: Even at 3-4 remaining, you'll still win at the next turn. Above that your chances of getting it right are too low compared to the reduction (assuming there is an option eliminating enough).
This could be made more complicated/interesting if you play a series of games and are awarded points based on either how many rounds it took to win or how many remaining cards you still had.
Oh fair enough. My apologies Mr Worf. I don't fully agree - plenty of shitty behavior gets ignored (or even encouraged) even in a workplace - but there's definitely some truth here.
reply