I judged three hackathons over about ten days this July: MLH x DigitalOcean “AI for Social Good” on July 11, the Sports World Cup Hackathon in San Francisco on July 17, and Aethera Hacks, an online event on Devpost, across July 19 to 21. I came in from the sports technology side, building athlete monetization tools, so I was usually the judge asking who pays for this rather than the judge asking what is your bundle size. That turned out to be a useful seat, because the questions that decide scores are mostly not technical ones.
Here is the part builders rarely get told: a judge is scoring under a hard constraint. Some number of teams, a fixed window, and by the middle of the block the demos start blurring together. Judges are not evaluating your project against an ideal. They are ranking it against the six they just saw while trying to remember which one had the map. Everything below follows from that.
The rubric is real, but it is not what separates teams
Most events hand judges four or five categories with numbers next to them. Technical difficulty, originality, design, impact, something about use of a sponsor API. Those categories are real and I filled them in honestly. But they compress. Almost every team lands mid-range on most of them, and the spread that produces a winner comes from two or three things the rubric does not name directly.
1. Whether the demo ran
This sounds too obvious to write down. It is the single largest score differentiator I saw. A working demo, live, on the judge’s screen or the team’s laptop, beats a more ambitious project shown as slides almost every time. Not because judges are impressed by working software as such, but because a live demo removes doubt, and doubt is what a judge is actually managing under time pressure.
The practical version: cut scope until something end-to-end runs. One complete path through the product beats four half-built paths. If your architecture diagram has six boxes and two of them work, demo the two and describe the rest in one sentence.
2. Whether you said what it is in the first sentence
The teams that scored well opened with a plain declarative: this is a tool that does X for Y. The teams that lost ground opened with context. Market size, a personal story, a problem statement that took forty seconds to arrive at the product. By the time the product appeared, I had spent a third of the slot without knowing what I was looking at, and I was reading the rest of the demo trying to catch up rather than trying to evaluate it.
Judges are not hostile to your story. They cannot hold it in memory without an anchor. Give the anchor first, then the story fits somewhere.
3. Whether the scope matches the time
A team that built one narrow thing well reads as a team that made decisions. A team that built a platform in 36 hours reads as a team that did not. I scored the narrow projects higher consistently, and so did every other judge I compared notes with. Ambition is not the signal. Judgment under a constraint is the signal, because the constraint is the only thing the event actually tests.
What I asked, and why
My questions were nearly always the same three, in some form.
- Who is this for, specifically? “Athletes” is not an answer. “A regional fighter with 40,000 followers and no manager” is an answer. Specificity tells me whether you talked to anyone or guessed.
- What part did you build this weekend? Not a trap. I want to know where the work went. Teams that answer this cleanly, including “we wired existing pieces together and the new part is this,” score better than teams that blur it. Blur reads as hiding.
- What breaks first at scale? The best answer is usually a specific known weakness said out loud. Teams that say nothing breaks lose credibility instantly, and teams that name the weakness gain more than they lose by naming it.
That last one is worth sitting with. Admitting a limitation raised scores in my sheets. It is counterintuitive if you think of a demo as a sales pitch. It makes sense if you think of a judge as someone trying to decide whether to trust you, which is closer to what is happening.
Things that did not move my score
- Polish beyond legibility. Clean, readable UI helps. Elaborate UI does not add on top of it, and past a point it makes me wonder where the backend time went.
- Framework choice. Nobody scored higher for a stack. A few teams spent demo time on stack decisions that could have gone to showing the product.
- Team size or credentials. I did not know or care who had which job. Solo builders did fine.
- Slide decks. They fill time. They do not persuade. The only slide I ever wanted was one screenshot of the thing running, as a fallback if the wifi died.
The thing judges never say out loud
Judges score partly on whether they can imagine explaining your project to someone else. That is the real memory test. When the panel reconvenes and someone says “which one was the scheduling one,” the projects that survive that conversation are the ones with a one-line identity. Projects with three features and no center do not survive it, even when they are technically stronger than the winners.
So the practical instruction is unromantic. Pick the sentence you want a judge to repeat to another judge an hour later. Build the demo that earns that sentence. Cut everything that does not.
If you are entering one this year
Five things, in the order I would do them.
- Decide the one-sentence identity before you write code, and check every scope decision against it.
- Get one end-to-end path running by the halfway mark, however ugly. Improve it after, do not extend it.
- Rehearse the opening thirty seconds out loud. Not the whole demo, the opening. That is where the score is set.
- Record a 60-second screen capture of the working flow as insurance. Wifi fails at most events I judged.
- Prepare an honest answer to what breaks first. Have it ready before a judge asks.
None of this is about being a better engineer. Most of the teams I judged were competent. The gap between the top table and the middle was almost entirely a gap in how the work was framed under time pressure, which is a skill the event tests whether or not it says so on the rubric.
I am the founder of Gameplan, a platform for professional athletes. I judged MLH x DigitalOcean “AI for Social Good” (July 11, 2026), the Sports World Cup Hackathon in San Francisco (July 17, 2026), and Aethera Hacks on Devpost (July 19 to 21, 2026). These are firsthand observations, not aggregate data, and no project, team or score details from any event are disclosed.
답글 남기기