Skip to content
SunnyWriteUps
Go back

The Ultimate AI Hackathon Evaluation Guide Judging Antigravity & AI Studio Projects

Edit page

Whether you are organizing a community hackathon or sitting on the judging panel, here is a comprehensive guide and checklist to ensure you evaluate AI projects fairly, rigorously, and effectively.

What is the Challenge About?

Before diving into the code, align the judging panel on the core objective of the hackathon. For a modern AI hackathon, the challenge typically revolves around:

Overall Guidelines & Instructions

When setting up the evaluation process, communicate these non-negotiable instructions to both judges and participants:


The Essential Evaluation Checklist

Use this checklist to vet every project before assigning a final score.

1. The Technical Baseline

2. Deployment (Additional Points)

3. Team Dynamics & Roles


Evaluation Criteria Rubric

Score each project out of 100 points based on the following breakdown:

CriteriaWeightDescription
Technical Depth35%How well does the project utilize Antigravity’s agentic features or AI Studio’s multimodality? Is the code clean, functional, and well-prompted?
Innovation & Creativity25%Does the project solve a problem in a novel way? Are they pushing the boundaries of what Gemini can do, or just building a standard chatbot?
Impact & Utility20%Does the application solve a real-world problem? Is there a clear target audience that benefits from this solution?
UX & Presentation10%Is the interface intuitive? Does it handle AI latency gracefully (e.g., loading states, streaming text)? Was the pitch compelling?
Completeness10%Is it deployed on Google Cloud? Is the agents.md thorough and accurate?

FAQs for Judges

Q: What if a team uses a different LLM in the backend instead of Gemini? Unless the hackathon rules explicitly allow multi-model architectures, projects must strictly adhere to using the Gemini API via AI Studio or Antigravity to qualify for the main prizes.

Q: How do we verify if they used static data instead of live AI generation? During the live Q&A, ask the team to input a highly specific, obscure, or contradictory prompt that they couldn’t have prepared for. If the app handles it dynamically, it’s live. If it breaks or returns a generic response, investigate their data flow.

Q: What exactly should be in the agents.md file? It should read like an architectural diagram for AI. It needs to list the agents used, their specific system instructions, the tools they have access to (like search or external APIs), and how multiple agents communicate with each other if applicable.

Q: A team has a brilliant idea but the Google Cloud deployment failed at the last minute. How do we score them? Evaluate the local build based on Technical Depth and Creativity, but deduct points from the “Completeness” category. A working local prototype demonstrating complex agentic behavior is always better than a broken cloud deployment.

Developer Journey and Operational Flow Registration and Capacity: The area operates on a first-come, first-served basis with no pre-registration. Each cohort holds a maximum of 25 participants, and attendance is logged via QR code scanning upon entry. The Coding Session: Each session lasts 50 minutes total, consisting of a 40-minute coding challenge and a 10-minute evaluation period.

Challenge Structure: Participants are guided by facilitators and have access to wired internet. Staff will provide subscriptions to participants using their personal (non-corporate) accounts. Participants are encouraged to share their work online after completion and “hit the alarm” at the exit as a celebration. Evaluation Area: Located on the area, evaluators—including Experts will assess submissions.

Evaluation Guidelines Mandatory Requirements: Usage of the Gemini API is mandatory, and solutions must demonstrate actual AI capabilities rather than relying on mock data. Verification Process: Evaluators should verify if the project solves the problem statement, confirm it works (locally is acceptable), and ask simple questions about the technology stack to ensure the code is not hardcoded. Efficiency: Evaluations should take 2–3 minutes to keep the process quick and engaging.

Staffing Responsibilities Support: Staff members are responsible for explaining the challenge, assisting with development queries, and managing slot queues. Operational Discipline: Sessions must start on time regardless of whether all 25 seats are filled. If a queue forms, staff should manage expectations for upcoming slots. Resources: laptops will be available for participants who do not bring their own laptops. Preparation: Staff should review the provided slides and challenge material, including testing Android app development via AI Studio, to create a handy FAQ guide.

Important Logistics Cloud Credentials: Workshop attendees may have additional cloud credentials that expire at the end of the day, which are distinct from the Vibe Coding activities. Rehearsal: A rehearsal for the staff should needed to familiarize the team with the space, technology, and movement plan.

these are the instructon Skills and different roles


Edit page
Share this post on:

Previous Post
Building Autonomous Agents with the Gemini API From Scratch to Scale
Next Post
Unlocking the Power of the Web A Guide to Native Browser APIs