Five Engineering Habits That Separate Working Demos From Reliable Software

  • Thread starter Thread starter Vaishnavi Khosla
  • Start date Start date
V

Vaishnavi Khosla

Guest
Good ideas are rarely the limiting factor. In reviewing dozens of AI hackathon projects, the pattern that would reliably predict success had nothing to do with how brilliant an idea it was. Five distinct engineering habits were responsible for success, all of them easy to implement and all of them missed by virtually every team in a rush. This isn't hackathon-specific either. These are the five habits that make the difference between software that works and software that you can rely on.

1. Document Your Failure Inputs Prior to Coding the Happy Path​


Almost every single time when you progress from "I tried this, and it worked" to "this will work", you will find yourself having to write a single edge case test. While I watched one project summarize the input documents, I realized that it was crashing right away after receiving an input document that differed from the project's prototype just slightly- some other encoding, nesting of additional sections, or other differences that are not unusual. There was a single assumption about the input format in the code, and there was no way during processing of it to verify this assumption.

What should you do in such cases? Before coding the main function, that actually does something useful, write down three inputs that you think will cause failure in the process and describe the way how the program should respond to them: should it generate an error, should it handle them gracefully or should it retry the same operation. If you cannot provide what will happen in these cases, then you did not test the assumption that you made.

2. Failure handling needs to be considered while planning a new feature, not when creating a separate ticket for it afterwards​


Validation, retry logic with a backoff, idempotency of writes, and the fallback logic that is used when the timeout happens are not present in the demo, but these are the features that set apart reliable software from the one created on hopes. Those teams that produced the best demos did not use any advanced models or frameworks. These are the teams that considered the failure scenarios, timeouts, partial responses, and downstream service downtime at the very beginning of scoping the feature and implemented at least one of these in their first iteration, instead of releasing the happy path and then "implementing error handling separately."

This usually does not happen neither in the course of 24 hours nor even in the following sprint, as there are other tickets that need to be implemented. In practice, this means that when creating acceptance criteria for the feature, you should always include at least one failure scenario. Exponential backoff with the correct timeout takes only a couple of lines of code.

3. Specify the life cycle of your data before you collect it​


"Where and how long will we store it and who will get access to it?" would be a natural answer to provide in one sentence before any feature starts collecting user data, but not nearly ever done until someone actually asks. I witnessed a development team build an awesome resume/job postings matching service, and when they were asked about how they stored uploaded resumes and what their retention policies are, there was even some hesitation. Not because of negligence, but just because they haven't gotten that far yet, given the limited amount of time.

And this applies naturally to production environments, where "we will think about retention policy later" becomes "we simply did not think about it." The actual trick here: before introducing any feature which collects and/or sends out user data, write one sentence which specifies where the data will be stored, for how long, and who will have access to it, and then delete the data through the same pull request in which its collection is introduced, not through a separate one. If you need to think about encryption of the data as well, do it right along with schema definition.

4. Gain validation through one external feedback before you say the job is done​


Those who had an external mentor or another teammate and asked, "Does this really make sense to you?" managed to build much more usable interfaces compared to those who built something for an imaginary user. This is not about design skills or additional time spent. The issue is that you cannot see the exact thing that might confuse a newcomer after spending hours analyzing your interface.

The practice is: Ask one person who didn't create it to try your interface before you say the job is done. It will take you only five minutes but will help you find the confusing step that you stopped seeing after hours.

5. Write code in a way so it could be easily understood by the engineer who will read it in six months' time​


Clever implementations did not win the contest in evaluating the quality of conversation. Clarity won. Teams that managed to concisely explain their architecture and point out the place in the code where the decision was made were considered more credible because clear code means understanding of the system by its author.

The test is: If you are not able to give an explanation of the purpose of a certain module and its implementation details, it is not done yet, even though it works for now. It is true for any hackathon project and production system; the only difference is the speed of realization of the price of such a shortcut.

None of the practices described above is related to AI in particular. They are the things that can be easily realized and even easier to stop if you are pressed for time. What judging the hackathon proved is that each shortcut mentioned above corresponds to a certain failure of production systems: the untested input data that fails at presentation time, the missing retry logic that is never implemented and data-related issues that no one knows the solution for.
 

Thread statistics

Created
Vaishnavi Khosla,
Replies
0
Views
3
Back
Top