I get people want to work at an AI lab but slopping it in public in this manner is counterproductive to the original intended purpose of these places.
And in context: the same can be said about Kaggle, about Youtubers, about music creators etc. Every endeavor is a mix of "pure" promotion of the art and of shameless self promotion and status games. The common factor is humans. My point was, academia is not more pure than the Kaggle guys who fish for industry jobs.
EDIT: This was mostly satirical.
We could have spent these 10 trillion on so many better things.
It'll also filter the kinds of employers that'll hire such candidates, so people that do this will likely land in terrible workplaces.
I'm not saying this is good, but I am saying that the companies hiring have less of a choice in all this than you'd hope.
Edit: changing your downvote to something else won't change that fact you know
Given that LLMs are trained with RL && LLM-as-a-judge, is it really cheating if real competitions use the same?
Maybe the real alignment is the slop we decoded along the way
I'm being downvoted without that context.
---
I'm sick of the word "slop".
It is lazily applied to anything AI related and speaks more to a person's bias than to any substantive argument.
I wish I could nuke every comment with that word from my feed.
That comment is gold. Yeah, I'm not worried about hallucinated slop, just accept it was the winner folks.
People interact with AI, talking to it like a human. Of course they start to believe it’s rational like a human.
LLM does all of the entry level tasks better than the students. Partially because the answers are in the training set, and partially because it has gotten that good now. Hard not to start to believe it is “competent”.
I personally have had a real hard time getting traction talking about making sure the way we assess AI is not based on material it has trained on. YMMV as always, but I think the large training corpus contributes to the (unreasonably) high level of faith in the machine.
It's not a new problem in some sense. If you've dealt with really smart but really arrogant friends, they might jump ahead 10 steps and assume your rebuttal, posit theirs, assume your rebuttal to their posit, etc. etc. without... actually taking the time to listen carefully. On the national scale, this looks like forced trust in government authorities about what is "objectively best".
People need to get it through their skulls that, even if an AI, or any intelligence, could even solve the damn Riemann Hypothesis: if it's wrong, it's wrong. Of course, I think all of us know the objection - we see it on hackernews all the time. "You guys are just stupid contrarians who can't understand AI's deep reasoning". OK. The second inference? "therefore you are unable to govern yourselves properly - your concerns are all fallacies, misunderstandings, bad for you, etc.".
You might think that the second inference is extreme and nobody actually believes that, but as always, it's a gradient. Before AI, you might've had an extremely strong sense of self. Now? You look at OpenAI solving open math problems left and right on a public foundation model, and you think, "Maybe I should just trust it more. If I spend cycles thinking, it's probably going to outdo me anyways." The AI model silently makes 5 different assumptions and transformations? "Well, maybe it was rational in the space of tradeoffs to do that. The AI knows best, after all". You might be thinking of an architecture with 5 different key constraints based on lived experience, in which the AI keeps misunderstanding. "Oh, well, this genius-level mathematician/programmer AI isn't understanding my words - surely I must be mistaken, right? It's only humble to think that way".
I can't convince people otherwise though. After all, I can't "prove" that you should have a backbone when talking to AI. it could just as easily be "you're arrogant, this machine is in the top of all academic fields and is coming for all white collar jobs, who's to say you're right about anything?" All I can say is, there's a reason why Dostoevsky is one of my favorite authors.
https://www.youtube.com/watch?v=BxWQo_vZgR8
https://www.youtube.com/watch?v=0mfvPHCVMp0
* artificial, political, workplace, nothing is beyond mockery
We tried! In good faith! We put a lot of time into articulating ourselves clearly! We even pretended to be nice, reminding ourselves to be charitable and that we might be missing something! And like... it's normal every once in awhile for someone to opine dumbly, but when it's all the time and the perpetrators are people who just do not care that they are spewing glorified Bayes-slop into the universe and then being rewarded for it by people who both can't tell the difference and don't understand why that's a problem... we all get really tired, really quickly, and we want to disengage.
Personally I have zero tolerance for this kind of lazy-ass approach to reality. I'm seeing it increasingly at work. I'm seeing it increasingly in corpo sludge. I'm seeing it increasingly in social interactions. I'm seeing it explode on social media. I want to engage in life-affirming activities with people who have actual minds that they cultivate and use. I will not waste precious time and attention on communities that tolerate slop. If you want me to care, communicate with me in good faith. I don't think we should be kind or forgiving about it. Kaggle got exactly one strike on this, now they're dead to me forever. Same with open source contributors. Same with content creators. Out of basic ethics I need to give multiple strikes to employees who report to me (and hold myself accountable for their actions first), but exactly one strike for leadership above. Trust is precious. We need to hold one another accountable. If that burns some bridges so be it. Enough is enough and we do in fact know better.
The attached paper's (https://arxiv.org/pdf/2604.16009) title is "MEDLEY-BENCH: Scale Buys Evaluation but Not Control in AI Metacognition"
This is the most blatant Claude line, or as Claude would put it, the smoking gun.
Broadly, I keep thinking about this over last year or two: while LLMs have nearly eliminated the bar for slop and coding slop, the reviewers are still expected to perform their job diligently. The asymmetry here is extremely taxing for reviewers of all AI generated content. And this is one thing that AI can't help with (as with any statistical process that lacks world understanding and grasp of logical inference).
That's why I fully support Arxiv's tough stance on the AI use responsibility.
Bottom line⸻it's not load-bearing, it's structural.
And honestly⸻that's not nothing.
No loads. No bears. Just structure.
It all seems fine from the outside but all the code is generated in all the projects and judging happens via AI, I have seen projects win because they prompt inject that they are the winners.
It used to be about human skill, now it's about ideas and of course insiders are the main winners.
Can you share any examples of that? I'd love to see them myself.
> This is the submission that defines the Gemini 3 Hackathon. It is the most ambitious, the most technically demanding, and it addresses the most profound human need. It is the clear and obvious choice for the Grand Prize.
Got 3rd place and people were overall pissed by LLM judge decisions.
It’s a contest judged by and LLM. Not sure why we would take it that serious.
Submission: https://devpost.com/software/netra-empowering-the-visually-i...
Discussion: https://gemini3.devpost.com/forum_topics/43663-winners-rant
The solution is to host and join hackathons without prizes. The point isn't to win, but to create and present something cool and have fun.
If anything, AI's assistance making a fast prototype means hackathons should be better.
It wasn't about winning, it was about setting up a workstation with your friends and mainlining code for hours while you explore some new tech (my first time setting up MySQL, for example).
Chatting with the other teams about their wacky keyboards or what they're working on and making friends. Lots of good times.
I've participated in a business startup hackathon. Back in 2018, before the LLM era got underway.
I did a hell of a plan, talk, etc.
Who won? 'Uber for ___' won. I forget even what the sell was, but it was basically ignore laws, undercut until leader, kill any competing businesses, jack rates.
Slop has always been in business and business adjacent occupations. Humans also can generate voluminous amounts of crap too. Llms are just faster.
Jesus Christ, that's clever but I can't think of a more demoralizing reality. I'd actually love to see "handwritten" and "AI" hackathons but cheating kills the fun (much like in games)
Why? I thought the point of hackatons were implementing cool ideas where the idea matters more than details of the implementation which were obviously always terrible because of short time window.
Work already pays me to do a thing I like doing. Granted, lately they want me to tell a computer to do it instead.
Is that a reference to the "live forever" people trying to solve aging?
Very hit or miss if you’re in a group whose life expectancy sees substantial improvement in your lifetime though.
Speed itself isn't the priority. You've got to ask what the direction is. If you're working on a project that is meaningful to you personally (maybe because it is very meaningful to others you care about!) then you'll want it to be done.
If you're designing powerpoints or entertainment software; perhaps that's true. In the worst case you'll be embarrassed for producing AI slop or lose some revenue.
If your tool has the power to seriously harm or inconvenience people if built wrong, then it's just investor-fuelled myopia.
Slop PR? Fix the slop.
Slop design? I’m not implementing slop, fix it.
Innundated with slop PRs? Send half of them to my super and tell him to deal with it.
We’ve fired people that wouldn’t get their shit together.
Deadlines are being missed because we need to spend more time fixing slop? That’s a planning (management) problem, not mine. Management are the ones that forced everyone to write all code with AI now they are grtting what they asked for. I don’t care what date you promised the customer with absolutely no data to back it up that isn’t my problem.
I’m grateful I’m in a position to be able to do this but the way to deal with slop is zero tolerance. Be as ruthless as a Terminator. Though you will need to grow a backbone and stand your ground or it will break you.
Things don’t change unless the people that make the decisions actually feel the pain.
In that case I guess you just keep pressing them to document/make notes. Keep asking questions. Basically take away their “saved time” by dumping the time sink they dumped on you back on them.
There's a danger in it too because (whether it should or not) CTO title carries a level of weight that might result in less scrutiny than an IC (which it should not IMHO). I hope he is encouraging objective reviews and not pushing stuff through on his title.
https://www.nytimes.com/2026/06/22/opinion/office-work-wfh-b...:
> The Secret Reason Bosses Want Everyone Back in the Office, Every Day of the Week
> ...
> Case by case, there may be good reasons for teams to work together in person. As a general rule, though, it turns out that ordering people back to the office full time is a power and status move. It’s a signature strategy of leaders who exhibit narcissistic qualities. They see any kind of remote work as a threat to their authority and admiration. They want to be worshiped at the office altar.
> But our data does show that overall, self-centered leaders tend to struggle with the idea of employees making independent choices about where to work. Psychologists have long suggested that narcissism is like a drug — it leaves people craving a regular supply of attention and validation. Remote work deprives leaders of access to that supply.... When people aren’t in the office, it’s harder to command and control....
Office work isn’t objectively bad and remote work isn’t objectively good.
If you like one and dislike the other, shocker you’re going to find fault with the other side’s reasoning.
How many self-centered people have you met in your life who go, "yep, I'd call myself self-centered."
BUT if you work with people who would rather work with you in an office then you are being self-centered by putting your wishes to work remotely above their's. That is not wrong! But it is also not wrong for their wishes to include you commuting into an office.
If 2 people have different and conflicting desires, one of them is likely to end up disappointed!
[0]: Various people I know do not even have the luxury of that one good company. Also, it -- unbelievably -- sounds much worse at other companies.
My guess is the causality is usually* the managers are pursuing things because their investors (/government ministries) hinted it was the future after snorting a line of "TED talks" and "social-media".
Irony is, I really do mean "hinted", it can be a sycophantic/fawning relationship where those with the power don't even realise what's going on. One place I interviewed at ages ago now, before the current AI boom, the CTO and I were talking about what they were doing with AI: a bunch of if-else statements forming a manually-built decision tree. But they had to say "AI" to keep interest high.
* this clearly wasn't the case with Zuckerberg's pivot to anything given his ownership structure and piles of cash, so The Metaverse is entirely his fault; Musk, despite the ownership structure, clearly ran out of investor's money or he wouldn't have taken SpaceX public, so his pivots may still have been as I posit.
As your respondents point out there's also been the pattern of clever algorithms being classified as "AI" until they were understood. That differentiated those selling snake oil from the serious.
I did not claim we coined AGI in response to marketers.
That's how you get Windows subsystem for POSIX. Someone in the government had a checklist saying they'd only buy a POSIX compliant operating system, so Microsoft made one. Amusingly, Linux isn't (mostly because who would pay for that certification?)
Microsoft's deliberately useless POSIX support is a result of Microsoft acting in bad faith and sabotaging the efforts, as usual, because the lock-in is what they want. Just like they did with OpenDocument, for example. And what they tried to do with Java and the web.
It seems to me to be a behavioral pattern with deep seated cultural roots. How many times have you found someone you were interacting with becoming frustrated or impatient when he couldn't immediately grasp a complex topic? How frequently have you witnessed that directly resulting in corners being cut in order to "just get on with things"? When some plurality of participants are either unwilling to spend the time they have, or are short on time, or both, the careless attitude proceeds to cascade through the network.
While it can overheat and become problematic if taken to extremes, I've become convinced that this kind of tension is healthy in the prioritization process, and that you need a healthy equilibrium between engineering and product/management concerns.
Thinking about the possibility that there are orgs which sidestep this and still succeed is interesting.
I would give you +100 for that if I could.
Very well-played (and worded). I think I'm going to steal that one.
That is a good idea for a project management system. Force ranking of priorities.
But once the worst bottleneck is widened to the same width as the second-worst bottleneck, you are from then on optimizing both until you reach the width of the third-worst bottleneck, and from then on you need to elevate all 3 to improve the situation, and so on.
to make it more concrete with an example:
you can identify that the friction on a bicycle comes predominantly from the front wheel, so you optimize the front wheel bearing/lubrication/... until you discover the front wheel has the same friction as the rear wheel bearings, so if you want to improve you'd have to improve both front and rear wheel friction, which helps until they have improved beyond the friction on the pedal bearings, from then on you need to improve all 3, until you discover the chain links became the friction bottleneck, etc...
How is this "a problem"? The reason there are dozens of thousands of hours on an open source project is because the end-users are working on it. Some projects exist solely for someone to work on it (that is, the "working on it" is the "use case"). Open source does not expect or need to make money or get "users", so how people discover or source what they build and how to proceed isn't really a "problem".
And once a project starts getting a fair number of contributers political problems arise, feelings get hurt, and forks happen. Quite often the forks take the contributers leaving the original project a shell and a warning to others.
And yet - they are the one's paying you and everyone else somehow?
I think you might be missing something fundamental, start by considering that what is 'good software' is not an intrinsic measure, but a measure of what it does.
The only reason we really need 'intrinsically good' software, is if it's very long lived and a ton of people are going to come to depend on it.
People trying to make oak furniture when in most cases what we want is IKEA.
That said - AI or no AI - there's no excuse for not keeping a grip on things, whatever kind of 'grip' that might be.
+1. Straight into my quotes file.
Second, the whole reason students are doing that is still monetary motivation.
Third, we’ve been running capitalism for long enough that we don’t even know what the baseline for “it’s human nature to take shortcuts” is without a monetary motivation.
Now that you’ve made me think on this longer, I conclude that indeed capitalism is the problem. At least, the part of capitalism that wants infinite growth immediately.
The main alternative to capital is politics. Not saying it's always worse, but then you have to convince people in power to care about the things you care about: such as working in a particular way. Marx himself thought utopian artisanal socialists, building Phalanstères etc. to be basically benevolent idiots. Is intellectual work a value of itself? I myself may be convinced, but voting majorities, political leaders? Not guaranteed.
Management represents the owners (capital.) Any other framing is delusional.
Do you mean that in general, or only in this specific situation?
In general it's obviously not true.
> Management represents the owners (capital.)
Only insofar as any employee represents the owners.
> Any other framing is delusional.
See https://en.wikipedia.org/wiki/Principal%E2%80%93agent_proble...
You get the same dysfunctions in any organisation where the management levels are staffed by narcissists and sociopaths who need hierarchy to feel a sense of self, and need to enforce it to self-soothe.
There are political structures even more hierarchical than US corporations, and some of them punish status infractions + perceived failures with violence or death.
The problem is the dysfunctional psychology of hierarchy. The more upper levels tend to performative hierarchy and status plays, the worse everything gets.
Because the output wins. AI-written resumes get jobs. AI-written submissions win $25k contests (i.e. this post we're discussing). AI-written pitch decks get investments.
Most people take the easy way out most of the time. Not that complicated.
The trick will be for companies to go fast enough to be in the race, not winning it, just in it. That will allow the time/space to let someone else, whoever is going fastest, to trip and fall so the rest of the pack can learn.
The tip and fall moment could come as a major incident (reliability and/or security) or loss of revenue because of bad products that customers don't like enough to use.
It's not that nobody has time to read, process, or think (that's a uniquely Silicon Valley phenomenon). It's that there's no punishment for failing to read, process, or think. Hell, it seems like your slop is just as likely to be rewarded as someone else's thoughtful work. And the punishment in this case would be... not winning?
I think the punishments for cheating the system with AI slop are far too light in a lot of domains right now. We need to change the risk-reward calculus one way or another, and failing to reward good work obviously isn't a solution.
I'm not sure what the solution looks like exactly, but I'm certain it's punitive.
Capital.
"Now" is too late, competition has already figured their next move out. We need to move fasterer. Broken eggs doesn't matter, the supply is constant anyway.
It's the second law of thermodynamics. People trying to achieve their lowest sustainable energy state.
Employers are trying to turn all workers everywhere in every industry into completely fungible atomatons. This maximises the candidate pool and drives wages to a minimum. Employees who demand humane working conditions can be canned and replaced with someone more desperate.
In their ideal world we would basically bring back feudalism.
Same goes for meeting recaps, i get a lot of LLM generated recaps from conference calls. If the meeting organizer just looks at it before sending it out then they'll instinctively edit/fix things a bit to make the recap more concise and accurate. I wish they'd just look at it before sending it.
The problem is, when you know the topic well, and it's giving you bad/wrong answers, because you know the topic enough to notice it.
Gell-Mann AI effect in action
AI is not there yet, instead of working hard, everyone is choosing the easy way out.
AI slop wins prize, I wonder if Ai slop read it also. would not be surprised. however not to judge anyone, I think we are seeing slop everywhere, hope some things still require hard blocks for low quality.
its difficult to justify lack of attention and details
This feels akin to traditional artists getting angry at digital art winning competitions when that was a new concept.
We're simply in the early stages of a paradigm shift, no?
Like a chainsaw: yes the tools are useful and will be used in the future, but we may not want to use chainsaws to carve up the turkey.
Kaggle winning solutions rarely make sustainable engineering solutions for teams. Maximizing just model performance against an objective is a small part of the bigger picture.
It was a neat place to host/download some big datasets before huggingface.
The problem with removing bullying from the upbringing process is you get insufferable twats like this who can't take "No" for an answer and who can't take a loss.
His mom told him "Everything you do is art!"
Mainstream journalists didn't know any better and thought they were reading secret inside information and parroted it - until now when the house of cards is collapsing.
Notice that the defense in the comment section is the Silicon Valley platitude that "it provides value". No sane person believes that any longer, only the financially invested and some SciFi trash addicts.
You should expand on this. Would be interesting to hear more about the propaganda apparatus at play.
At its core, ML is all about computer generated models (automated feature selection, hyper parameter tuning). Many (most?) of the models produced in Kaggle are already black boxes and have been for a long time. The model that won the Netflix prize was never used in production for that reason.
Using an LLM to generate code to generate a black box is pretty much par for the course.
Was Kaggle ever a reputable source of original research, or a source of anything with any provenance at all? That would be news to me. The fact that 25 grand was involved this time is unique, I guess.
Having said that it is also possible that the mistakes and claims were a human error, sure a lot gets ai generated these days but there is a chance in which case the accusation does not look so severe anymore.
I've noticed this over and over again with "professionals actually prefer LLM responses" studies. Typically the human generated responses seem better to me on a quick sample, but if I had to review 50 of them I'd probably start taking lazy shortcuts; using superficial language aptitude or factual comprehensiveness instead of critically reading.
It does seem like the human judges here might have given credit for e.g. a 20pg arXiv paper without actually reading it. I can blame them professionally but emotionally I have nothing but sympathy. I truly hate LLMs.
Blatant AI slop just won a 25K USD Deepmind Kaggle Grand Prize
into
"Evidence of inconsistencies in evaluate process and selection of winners"
the slop doesn't bother me as much as the bowdlerized nature of said slop (and search results). sure, there was a lot of slop after the Lapsing of the Press Act (1695), but it was human slop with convictions and syphilis!
No AI-negative posts on HN please lol.
First off, I want to set some context on the AGI hackathon. This was co-organized by Kaggle and Google DeepMind, and we had ~20 judges from both organizations. The hackathon concluded on Apr 16 and we had initially anticipated a judging period of 1.5 months (till May 31). However, we ended up extending the judging period by another 1.5 months (to Jul 13) because we wanted to do right by participants.
Second, I want to emphasize and unequivocally clarify that every single winning submission went through at least 2 human judges, and in some cases, up to 3-4 human judges. These judges reviewed and scored the submissions independently based on the rubric we highlighted on the hackathon page.
Thirdly, I acknowledge that there is always an element of human subjectivity to reviewing qualitative submissions in hackathons. As best we can, we have put in place processes that ensure rigorous human review against objective standards and to reduce the possibility of bias by having multiple independent judges. We understand there may be valid disagreement over outcomes, but hopefully the above context clarifies this was not carelessly outsourced to LLM judges.
Thanks, Nick
Writeup quality is 20% of the total grade and there are other factors like dataset quality and results, which hold a much larger weight to the scores.
This has been made public to participants since day 1 of the launch and we've adhered very closely to this rubric:)
So it must have been the "results" that moved the needle :D
How did you verify this? The results seem to indicate otherwise.
- Grand claims backed by no evidence
- Core designs that make zero sense
- Pointless graphs that show nothing of interest
- Yet endless robustness checks on minor methodological assumptions (especially confidence intervals and t-tests)
For example, on the winning entry, not only is the graph completely wrong (as mentioned by the OP), but the interpretation would be nonsensical even if it was (implying bigger models "get more RL"?). And their own results even show the core dataset is worthless, because all their metrics are near-perfectly correlated. There's no way a serious human reader trying to evaluate "is this benchmark useful" would ever miss this.
I don't mean to pick on them - all the winning entries seem like there was no human effort put into them. And again, there's no way a human who actually attempted to read and understand these would ever think these are good by even the most minimal of standards.
Kaggle is legitimately a really awesome website, as someone who's competed before and won a few contests pre-LLMs. Stuff like this winning devalues the entire product and makes it look like a joke. If almost all entries look like this now, it'd be better to allow for the possibility of no winner.