Getting AI into production: What we learned from our internal hackathon

How our Innovation Summit turns engineers' AI experiments into features clients now run in production.
You can't experiment on a product with live users and delivery deadlines. But if you don't experiment, AI stays on the roadmap.
We face the same problem. So we ran our first AI Innovation Summit, an internal hackathon where 48 engineers in 16 teams built AI solutions from the ground up. Several of those solutions are now part of real client products.
The Summit showed us what AI experimentation looks like in a service environment, where client commitments and delivery deadlines don’t pause for a hackathon. Below, we share what worked and where we stumbled, along with what it all means if you’re planning to bring AI into your product.
TL;DR | |
What happened | 48 engineers split into 16 teams and spent 4 weeks building AI solutions from scratch at the first AI Innovation Summit at VITech. |
Real production use | Several solutions are already running in live healthcare products, and more are on their way to production. |
Why it matters | AI ideas are built, tested, and judged on business impact before they ever touch your product. |
Why we launched an internal AI hackathon
In Q1 2025, we launched a Research and Innovation department that brought together our CTO, CDO, Head of Architecture, and Director of Engineering. Its first job was to make AI part of our teams’ daily work.
We started with training in machine learning, deep learning, and large language models (LLMs). From there, we moved into more practical topics: prompt engineering, agentic approaches, frameworks like LangChain and LangGraph. It was a good reminder of how fast this field moves. Prompt engineering had only just been hailed as a “new profession,” and by the time we got to it, the buzz was already fading.
Training alone, though, wasn't enough. Our engineers needed hands-on experience, and clients were asking about AI more and more often. Yet we couldn’t use live projects as a testing ground. When you’re responsible for a client’s product and its delivery deadlines, experimenting on it is off the table.
We needed something in between, and that’s how the idea of an internal hackathon was born. It was the first time we’d tried anything like it, so frankly, we had plenty of doubts.
Designing a format that fits around client work
We couldn’t simply put work on hold for a few days. We built the hackathon format from scratch and refined it as we went.
To avoid boxing teams into predefined ideas, we offered a couple of tracks:
Track A: Client projects. Teams could finally take on problems they’d wanted to solve for a long time.
Track B: Internal processes. Teams built AI tools to make everyday work at VITech easier.
We planned for 2 weeks, which felt like enough time to build an MVP without burning out. In reality, the hackathon took twice as long. Coordinating 48 people across different client projects was harder than we expected, and we had to push the dates back more than once.
The extra time had an upside: teams went well beyond MVPs, and some solutions looked nearly production-ready. But what should have been a focused sprint turned into a marathon on top of everyone’s regular workload. That’s a trade-off we weren’t willing to repeat.

What we underestimated at the start
Beyond format and length, once the teams got going, a few more gaps showed up.
Communication
We set up a dedicated Slack channel and assumed it would be enough. Some teams used it actively to ask questions and share progress. Others ended up working in almost complete isolation, and we only saw their ideas at the final demo, when it was too late to change anything.
Looking back, we were missing a proper kickoff and at least one mid-point sync. Both would have helped us identify problems earlier and give teams the support they needed.
Mentorship
We initially saw mentors as a support system. In practice, some teams drifted toward ideas that didn't solve a specific business problem or fit the judging criteria. Mentors could have steered them back on course during the first week.
Judging criteria
From the start, we wanted clear, transparent rules, so teams would know we didn’t choose winners based on their luck or the prettiest demo. We judged on:
Innovation and creativity (30%): How original was the team's approach? AI had to be at the core of the solution, not simply added on top. We looked at both the idea’s originality and the depth of AI integration.
Business impact (30%): Could the solution improve a client product or an internal process? Teams had to explain where it could deliver measurable results and what should happen next.
Technical execution (20%): The solution had to run and do what the demo showed. A proof of concept or MVP was perfectly fine. We weren't expecting production-ready software, but we did expect working code.
Demo and presentation quality (20%): In 5 minutes, judges and the audience needed to understand the problem, how AI helped solve it, and why the idea deserved further development.
We kept fine-tuning the details and weighting of these criteria while the hackathon was already underway. As a result, some teams didn’t realize that innovation and business impact outweighed a polished demo. The right solutions still won in the end, but locking everything in on day one would have given every team a clearer target.

From hackathon projects to client products
Each of the 16 teams delivered a solution. Several went on to real client projects, and 3 are already in production. We can’t name the clients because of NDAs, but we can share the technical details.
1. Clinical Protocol Assistant
One of the strongest entries tackled a problem that directly affects how fast clinical trials can launch and how many errors slip through along the way.
The problem. Clinical protocols are complex documents, often hundreds of pages long, covering dozens of scenarios, forms, and data points. Preparing them for a client's healthcare product takes a specialist 4 to 5 full working days. Any mistake can trigger a protocol amendment, which means another round of reviews and approvals that hold up the whole study.
The team's solution. They built a pipeline that parses a protocol PDF and turns it into a structured configuration for the client’s product. A built-in review step lets a specialist check and adjust the output.
Under the hood, they used LangChain and LangGraph to orchestrate the workflow and Claude via AWS Bedrock as the model. One decision stands out: the team deliberately skipped RAG, or retrieval-augmented generation, the common approach where a model pulls in only the passages that match a query. As one team lead put it, RAG works when you know what you’re looking for, and with a protocol, you don’t know in advance which pages hold the key information. So the team summarizes every page in parallel and then builds up the protocol’s structure step by step.
The result. What used to take days now takes minutes, with a specialist still reviewing every output. After the hackathon, we presented the solution to the client, and it’s now on its way to production as part of their product.

2. Clinical Assistant Chat
The problem. Another healthcare client works with a large volume of clinical quality metrics. The data is there, but the interface for exploring it was limited. Users could filter and browse tables. But they couldn't ask something like, "Which patients need intervention most urgently across several metrics at once?"
The team's solution. The team added a chat layer on top of the existing data. A user asks a question in plain language; the model writes a SQL query to pull the right data, and the chat returns an answer with an explanation.
The stack included Python with FastAPI, Postgres, Claude via AWS Bedrock, and LangChain for orchestration. Responses stream in real time, and chat history is saved to a database, so users can keep several conversations going at once.
The team also put real thought into transparency. Technical users can see the generated SQL, while business users get a plain-language explanation. On top of that, automated prompt tests let the team keep improving answer quality without breaking what already worked.
The result. The solution is integrated into the client's environment and in active use.
3. Patient-facing Knowledge Chatbot
The problem. This client has a large library of content for end users, including articles, documents, and process explanations. Users struggled to find specific answers, so they either had to read through long documents or contact support.
The team's solution. Here, the answer was a RAG-based chatbot built into the client’s product, which lets users ask questions in natural language and get direct answers.
The team took a rigorous approach to the technical implementation. First, they built their own evaluation framework that combines keyword matching and LLM-as-judge, a method where a second model scores the chatbot’s answers.
With that in place, they tested several models, including Mistral and Gemini, and picked the one with the best balance of cost and quality. Embeddings also update automatically, so whenever content changes in the CMS, the chatbot’s knowledge updates with it.
The result. The client approved the implementation right after the demo, and the chatbot made it to production in weeks.
How AI is changing the pace of development
One of the most memorable moments came during the Q&A after the demos. A team presented an internal app built in 2 weeks. It featured time tracking, AI-assisted planning poker, semantic ticket search, profile management, and peer recognition.
On a client project, building something like that would normally take several months, so of course someone asked how they'd pulled it off. The answer: AI coding assistants had written part of the code. The app still needed refactoring, and some of the processes weren’t in place yet. But the structure and working logic came together remarkably fast.
That moment captured a bigger shift in software engineering. Writing code is becoming the easier part of an engineer’s job, and the real skill now lies in framing the problem well and making sound decisions about what to build. For your product, that means the bottleneck moves from writing code to deciding what's worth building. That's the judgment you want on your team.

What didn't work
Not every team made it to production conversations. Some solutions stayed at the MVP stage, which is perfectly fine. Not every experiment needs to become a product.
After the demos, some teams were left in limbo, unsure what to do next. We picked up the strongest solutions and deliberately left the weakest ones behind. But a middle group of projects was left without a clear next step. Those teams didn’t know whether to keep going or move on, and that uncertainty took a real toll on their motivation.
What we changed
After the first AI Innovation Summit, we learned from our mistakes. We went back to 2 weeks instead of 4, locked in the judging criteria at the announcement stage, and added mandatory sync-ups throughout. Post-hackathon follow-up is now built in as its own stage. This time, we're staying connected with every team to support them and answer questions as needed.
One of this year’s focus areas is automating internal processes with AI coding tools. It picks up right where use cases like that two-person team’s app left off.
Takeaways
The biggest change we saw was in mindset: teams stopped being afraid to experiment with AI. The risks didn’t shrink. We simply created a space where those risks weren’t critical.
Hackathon ideas made it into real projects, and in some cases, the path to production took weeks instead of months. Clients picked these solutions up on their own once they saw the business value for themselves.
None of this happened by accident. Looking back, a few things made this possible.
Environment. The Innovation Summit removed the biggest barrier, which is the fear of breaking something in production. Teams could try new approaches, make mistakes, and learn without touching client projects, and many took on ideas they’d never have found the time to explore otherwise.
Culture. We'd spent a long time building a culture of learning inside the company through internal talks, webinars, and knowledge-sharing channels. So the hackathon didn't come out of nowhere. It was a natural extension of that. We didn't have to figure out how to motivate people to sign up for such an extended extracurricular activity. Everyone was curious to give it a try.
Management involvement. Our C-level leaders took an active part. They provided resources and joined the reviews, which showed every team that this work mattered.
One more thing made a difference. When people know they’re building something that could reach production, their motivation changes completely.
What this means for your product. If you're planning to bring AI into your product, the same lessons apply. Test ideas away from your live product, judge them on business impact rather than demo polish, and plan the path to production before the demo, not after.
FAQ
How do you decide which AI prototypes are worth taking to production?
We start with the business problem. A solution moves forward when it solves a real problem in a measurable way and fits the client's architecture and security requirements.
How quickly can an AI prototype become part of a real product?
It depends on the project. Our patient-facing chatbot reached production within weeks. Most MVPs need more time for security hardening, automated tests, and data privacy checks before real users can rely on them. For healthcare products, that includes the same HIPAA-compliant process we follow on every build.
How do you turn a hackathon prototype into a production-ready solution?
We treat the prototype as a working blueprint. We review the architecture, check latency and API costs, and set up monitoring with a fallback before the feature goes through our usual QA process.
How can an AI hackathon help a company evaluate an idea before investing in development?
It gives decision-makers something real to judge. Our client saw the Clinical Protocol Assistant cut days of work down to minutes, and that's what moved the project toward production.
Share post




