Rendered at 23:27:23 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
edot 23 hours ago [-]
Tom Dietterich (Editor in Chief at arXiv) posted on LinkedIn the other day that they’re having trouble keeping up with the onslaught of AI-generated papers. Lots of suggestions in the comments but no magic bullets.
It’s ironic that the LLMs which benefit so much from reading arXiv papers of yore are now being used to pollute it. Personally if I see a single author post 2023, I assume it’s junk, especially if they’re not from an actual research institution. Not all solo independent researchers are phonies but … many phonies are solo independent researchers.
OpenReview is okay but even some of the reviewers are apparently using LLMs or just hardly reading.
bobmarleybiceps 23 hours ago [-]
A recent review of mine had a LLM-ism at the very end, "would you like me to format this into a formal peer review report?" So they very likely copy-pasted their whole review :). I'm pretty down on academia atm ;-;
Skyy93 22 hours ago [-]
I think this is more of a systematic issue. I review now since 2-3 years, I do not get paid, which is fine. However, it takes always a huge amount of time without really having anything from it, but I do it because it is important work.
There now so many researcher that need to publish which explains the flooding, LLM only speed it up, so reciprocal reviews take place. So now you are forced to review and you are having less and less time. So it’s a natural choice for you if you already took an LLM to write a paper to use it to review.
Perhaps one solution would be much harder entry barriers, and enforcing some guidelines. For example that a supervisor can not have more than 5 papers and PhD students only need one real paper on a major conference/journal.
coolness 16 hours ago [-]
Super unfortunate. On the other side, I reviewed four papers for a top-tier AI conference and three were clearly fully Claude generated, as in all text, figures, results, everything. Actual good reviewer time is wasted on such papers and your (i hope) human written good paper receives AI responses. It's a sad state of affairs for sure.
stingraycharles 11 hours ago [-]
I find it incredibly sad that this is what the world is becoming, and I mostly blame people for this, not the LLMs.
It’s the same type of people that would have no issue letting an LLM open a pull request on GitHub wasting valuable time of other humans, and whatnot.
I’m using LLMs all the time myself, but it’s so incredibly important to use it to improve the quality of your work, not degrade it. People seem to be totally oblivious about this.
On the flip side, it does make it easier to recognize people who are wasting my time.
jamienk 7 hours ago [-]
I think we underestimate the psychology of this kind of use of LLMs. I don't think these people (students, academics, lawyers, etc etc etc) are all just lazy morons. I think that using an LLM gives a strong knee-jerk feeling of "OMG this is exactly what I have to say. This is MY idea, these are my thoughts, this is - in a real sense - MY writing!" The feeling floods you when you see the output being churned out, way before you you read the thing (if you ever do) - it's the initial *seeing* it. Then when you hand the generated text in, unread by you, you do not have the feeling of cheating, you have the feeling of having exercised your powers, pushed right to the edge of your expertise and thoughtfulness, you successfully overcame obstacles because of your experience and unique capabilities.
This is not always the psychological situation, but I think it might be a lot of the time.
I also think that until we recognize this, we won't be able to help people to not do it. Calling them lazy or being bewildered by them or feeling rage or contempt towards them or threatening them is not going to be practically helpful.
baq 18 hours ago [-]
You’d better not look at the average CI pipeline in software shops then
Oh my god that was just written at the beginning of August?? It feels like I read it a year ago. We really are speed running this tech cycle…
htrp 7 hours ago [-]
the AI reaearchers who got their careers started on arxiv should chip in to support the site
auggierose 17 hours ago [-]
If you see any post on arXiv you should assume it is junk.
I find it ridiculous that people put any value on something being posted on arXiv. That doesn't mean the post is bad. It just means you need to find other means of judging it, for example by actually reading it.
gus_massa 12 hours ago [-]
I half agree. It used to be good, they called them "preprints" because they were already sent to a journal and somewhat expected to be accepted. Until people noticed that they were not forced to publish the "preprint" later so it got flooded with crap, hand crafted artisanal crap.
Unless you are working in the area of the paper AND know the reputation of the authors AND take a deep look, just give it the same credibility than to a random PDF posted in WordPress. They have some weak filtering because to post in the arXiv someone must vouch for you or something similar, but it's a very weak filter and people was already abusing it.
And then the AI slop truck hit...
jamienk 7 hours ago [-]
Did we completely write off the "web of trust" as a workable system? Where you exchange keys with people and co-sign each others keys and then trust keys signed by them. This would be an alt to "algorithms" and also a way to "subscribe" to things that got vetted by trusted people or orgs. I kind of do this with uBlock Origin, where they curate a (very complex) list of ads (but without the key signing). Orgs could vet and publish lists of arXiv papers and I'd be subscribed to that. This could also be the source of my social media feed (friends of friends, etc).
I remember in the early days of PGP this was put forward as a vision. I remember the various key-holding systems (that were hard to use). I did a key-party with my brother and his friend in the '90s.
Why did this fail?
minraws 6 hours ago [-]
Because then it would be a big club, but most wouldn't be in it.
To clarify: for a large number of topic I find this completely fair and valuable to first have to build trust. But getting your research paper in front of other humans on a aggregation website seems like the one place where this for me falls apart.
Today any one from any corner of the world from any educational or non education place publish stuff, well moderation is tight and often it won't get through but I know a few friends who aren't in academia that have a paper up there who wouldn't have been as easily able to get a paper on anywhere else.
Maybe that's not a genuine concern but that often used to be the case and still is in a lot of invite first clubs/groups. I know credible people in the field who were and are working in relative isolation which leads to not being able to participate in sharing of ideas.
aurareturn 13 hours ago [-]
Are they doing to do something about authors using Arxiv to publish propaganda/opinion pieces but presented as research?
ArXiv gives the appearance of scientific credibility that a blog post wouldn't have so I'm seeing the platform get abused.
Probably not because the whole point is they don't do critical review
> ArXiv gives the appearance of scientific credibility that a blog post wouldn't have
That's an error on your side not theirs
aurareturn 10 hours ago [-]
Then just call it Substack.
whateverboat 8 hours ago [-]
This comment is one of the most bizzare things I've seen here regarding ArXiv.
networkOne 21 hours ago [-]
Sorely, sorely needed.
How is science to evolve if good research requires $49 a pop to view?
And what amazes me is that these authors have undeniably stood on the shoulders of giants in order to create their research.
kadoban 21 hours ago [-]
You think authors are the ones driving the $49 view fees? They are not.
SonicTheSith 12 hours ago [-]
They are, if their institution does not have contracts for that specific domain from Elsevier et al.
At my institute we for example do not have access to parts of Springer.
Luckily, more and more is moving towards open access.
locknitpicker 7 hours ago [-]
> How is science to evolve if good research requires $49 a pop to view?
This kind of criticism is misplaced. The answer to reputable journals being held hostage by the likes of Elsevier is not dropping peer review and dump articles in a blog-like site such as Arxiv.
I'm confident the bulk of academia would love to move away from Elsevier et al and have free open journals as their discourse venue. However, the publish or perish situation created pressure to count only the papers published in reference publications in their productivity metrics, and those were hijacked by for-profit editorial companies which unscrupulously abuse their dominant entrenched position.
And publishing on Arxiv does not change that.
perdy 16 hours ago [-]
Published there in August, from a company rather than a university. Without arXiv we'd have had a blog post and nothing citable.
flexagoon 12 hours ago [-]
A blog post is just as "citable" as an arXiv preprint. There's no fundamental difference between them.
emil-lp 11 hours ago [-]
One is immutable with a DOI.
frumiousirc 12 hours ago [-]
> Published
Posted. Having a preprint on arxiv is not publishing.
perdy 3 hours ago [-]
[dead]
locknitpicker 7 hours ago [-]
> Without arXiv we'd have had a blog post and nothing citable.
Arxiv is a blog. Your paper is a blog post which is just as citable and as peer reviewed as blogspot.
liberian 13 hours ago [-]
[dead]
upupupandaway 23 hours ago [-]
Asking earnestly: is ArXiv a valuable resource? I used it few years back when finishing my (late) Master's, but also saw a lot of crap published there (primarily for promotion/visa purposes) so it kind of took the shine out of the service for me.
lemontheme 16 hours ago [-]
My take: if not for arxiv and huggingface, the field of ML would be nowhere near where it is today.
More to your question, I recommend something like semanticscholar to find actual relevant papers. Try to identify researchers that seem trustworthy, then explore the citation network around them.
For more hot off the press stuff, follow what gets boosted on social media.
Still doesn’t cover the truly niche stuff but it’s a start
txhwind 20 hours ago [-]
A free-to-read PDF host site is valuable enough. Most academic publishers have a pay wall.
tristanj 22 hours ago [-]
The preprint papers on arXiv are like 99% the same as the published versions, except they're free instead of locked behind a multi-thousand dollar/year paywall.
upbeat_general 21 hours ago [-]
And if there are differences, that is often a good thing! It can mean the author wanted to format something in a particular way that the journal didn't allow.
senderista 19 hours ago [-]
Often they are extended versions of the journal articles and contain valuable material that had to be cut for space limits.
locknitpicker 7 hours ago [-]
> The preprint papers on arXiv are like 99% the same as the published versions, except they're free instead of locked behind a multi-thousand dollar/year paywall.
I don't think you fully understand the problem. It doesn't matter if you can find in arxiv a preprint of an article published on a reputable journal. What matters is that right besides that paper you will find a dozen other papers that can be utter nonsense generated by a poorly calibrated slop factory. You don't find those in papers published in respectable journals which enforce double blind peer review and were filtered for relevance and quality.
It's that peer review process that creates value and relevance. Otherwise all you have is a glorified file server.
ufo 23 hours ago [-]
Some researchers and journals publish peer-reviewed work on arxiv; it serves as a stable archive that won't be plagued by link rot or paywalls.
edot 23 hours ago [-]
It’s good for getting free access to preprints which are often close enough to the paywalled real papers in journals. If I find a paper I want (or more realistically, if ChatGPT finds a paper it wants for me), but it’s behind a paywall, odds are the authors put a preprint on arXiv.
embedding-shape 22 hours ago [-]
Either you setup feeds for the specific topics/subjects you care about, scan what you come across once a week, or you use it to get access to papers that are usually behind some paywall. I don't think the intention nor the value comes from just haphazardously reading through everything in some section.
It's not peer-reviewed and supposed to free and accessible from both sides so the results kind of makes sense.
It’s ironic that the LLMs which benefit so much from reading arXiv papers of yore are now being used to pollute it. Personally if I see a single author post 2023, I assume it’s junk, especially if they’re not from an actual research institution. Not all solo independent researchers are phonies but … many phonies are solo independent researchers.
OpenReview is okay but even some of the reviewers are apparently using LLMs or just hardly reading.
There now so many researcher that need to publish which explains the flooding, LLM only speed it up, so reciprocal reviews take place. So now you are forced to review and you are having less and less time. So it’s a natural choice for you if you already took an LLM to write a paper to use it to review.
Perhaps one solution would be much harder entry barriers, and enforcing some guidelines. For example that a supervisor can not have more than 5 papers and PhD students only need one real paper on a major conference/journal.
It’s the same type of people that would have no issue letting an LLM open a pull request on GitHub wasting valuable time of other humans, and whatnot.
I’m using LLMs all the time myself, but it’s so incredibly important to use it to improve the quality of your work, not degrade it. People seem to be totally oblivious about this.
On the flip side, it does make it easier to recognize people who are wasting my time.
This is not always the psychological situation, but I think it might be a lot of the time.
I also think that until we recognize this, we won't be able to help people to not do it. Calling them lazy or being bewildered by them or feeling rage or contempt towards them or threatening them is not going to be practically helpful.
Unless you are working in the area of the paper AND know the reputation of the authors AND take a deep look, just give it the same credibility than to a random PDF posted in WordPress. They have some weak filtering because to post in the arXiv someone must vouch for you or something similar, but it's a very weak filter and people was already abusing it.
And then the AI slop truck hit...
I remember in the early days of PGP this was put forward as a vision. I remember the various key-holding systems (that were hard to use). I did a key-party with my brother and his friend in the '90s.
Why did this fail?
To clarify: for a large number of topic I find this completely fair and valuable to first have to build trust. But getting your research paper in front of other humans on a aggregation website seems like the one place where this for me falls apart.
Today any one from any corner of the world from any educational or non education place publish stuff, well moderation is tight and often it won't get through but I know a few friends who aren't in academia that have a paper up there who wouldn't have been as easily able to get a paper on anywhere else.
Maybe that's not a genuine concern but that often used to be the case and still is in a lot of invite first clubs/groups. I know credible people in the field who were and are working in relative isolation which leads to not being able to participate in sharing of ideas.
ArXiv gives the appearance of scientific credibility that a blog post wouldn't have so I'm seeing the platform get abused.
Example: https://news.ycombinator.com/item?id=49580164
> ArXiv gives the appearance of scientific credibility that a blog post wouldn't have
That's an error on your side not theirs
How is science to evolve if good research requires $49 a pop to view?
And what amazes me is that these authors have undeniably stood on the shoulders of giants in order to create their research.
At my institute we for example do not have access to parts of Springer.
Luckily, more and more is moving towards open access.
This kind of criticism is misplaced. The answer to reputable journals being held hostage by the likes of Elsevier is not dropping peer review and dump articles in a blog-like site such as Arxiv.
I'm confident the bulk of academia would love to move away from Elsevier et al and have free open journals as their discourse venue. However, the publish or perish situation created pressure to count only the papers published in reference publications in their productivity metrics, and those were hijacked by for-profit editorial companies which unscrupulously abuse their dominant entrenched position.
And publishing on Arxiv does not change that.
Posted. Having a preprint on arxiv is not publishing.
Arxiv is a blog. Your paper is a blog post which is just as citable and as peer reviewed as blogspot.
More to your question, I recommend something like semanticscholar to find actual relevant papers. Try to identify researchers that seem trustworthy, then explore the citation network around them.
For more hot off the press stuff, follow what gets boosted on social media.
Still doesn’t cover the truly niche stuff but it’s a start
I don't think you fully understand the problem. It doesn't matter if you can find in arxiv a preprint of an article published on a reputable journal. What matters is that right besides that paper you will find a dozen other papers that can be utter nonsense generated by a poorly calibrated slop factory. You don't find those in papers published in respectable journals which enforce double blind peer review and were filtered for relevance and quality.
It's that peer review process that creates value and relevance. Otherwise all you have is a glorified file server.
It's not peer-reviewed and supposed to free and accessible from both sides so the results kind of makes sense.