Dhruv Pathak, founding member of the Goibibo tech team, former SVP at MakeMyTrip, and co founder and CTO of INDmoney, joins host Kevin on Unclouded: The Agentic AI FinOps Podcast by Cloudgov.ai. Dhruv talks through absorbing 40% traffic spikes at market open, why travel tech and fintech engineer latency completely differently, treating compliance as the price of “peaceful sleep,” and why the real cloud bloat is human induced inefficiency, not security. Connect with Dhruv on LinkedIn.
Full transcript
Chapter 1 — Intro: From Goibibo to the INDmoney SuperApp
Kevin: Hello and welcome back to Well Operated, the podcast where Agentic AI turns FinOps into autonomous action. My name is Kevin. I am the human in the loop. Thanks for joining me today. Also with me on the show is Dhruv Pathak, a founding member at the Goibibo tech team, a former SVP at MakeMyTrip, and he was the co founder and CTO of INDmoney. As a Cloudgov.ai partner, he’s moving beyond the era of static dashboards and manual spreadsheets. We’re going to talk about how he balances massive market spikes with strict regulatory guardrails. Dhruv, welcome.
Dhruv: Hi, Kevin. How are you? How’s the new year going for you?
Kevin: Yeah, it’s going good. It’s going good. I can’t complain. I’m excited to chat with you about this. You’ve had a very fascinating career and I’d like to dive right in. You were managing an environment where traffic can spike by 40% in minutes during market opens. In a world where millisecond level trading is the product, how do you prevent your cloud bill from spiraling out of control during those surges?
Chapter 2 — Handling 40% Traffic Spikes During Market Opens
Dhruv: Yeah. So the good part is, I’ll talk about INDmoney context first. The fintech where one major use case for which we have our infra is trading, and trading is bounded by specific patterns. So you have market open times, then you might have certain volatility in the market due to some global news or let’s say some domestic news. Then you have a market closing time, then you have days where the market does not work at all. It could be national holidays. It could be festivals or weekends. And similarly we also have a market leading US stocks product. In that too the similar patterns follows.
So we try to put our scaling policies in such a way that they are always cost optimized based on these traffic patterns, and they also ensure that the end user who is putting money at stake, when he is pressing a button on our app, that user gets the most reliable trading or let’s say charting experience when they are doing any activity on our app. So it’s a balance between both the things. And the second point is of course compliance and security which we can dive into detail later on. So compliance and security also increases your cost, but it also provides you long term continuity of the business. So the cost optimization there is done with the balance of ensuring that you are compliant and secure, and while doing your FinOps activity within that scope.
Kevin: Okay. And looking back, do you feel like your approach to infrastructure has changed from your time building travel tech at Goibibo to managing the sort of high stakes, high frequency world of fintech?
Chapter 3 — Travel vs. Fintech: Accuracy and the 20ms Threshold
Dhruv: So the basic tenets and the basic DNA stays the same. You always have to respect your customer for their lifetime on your app or on your platform because they are trusting you with their hard earned money. Kevin right now is doing a podcast. With this hard work, he’ll earn some money, and let’s say he was a user of Goibibo and MakeMyTrip. He will use that hard earned money to book a flight or a hotel. And let’s say Kevin is an INDmoney user, he will use that hard earned money to maybe put a trade or do long term investing in a good tech stock.
So as a platform it’s our responsibility we ensure Kevin’s money is safe and Kevin is reliably able to do what he wants to do, time after time, every month, every day, whenever he wants to. But there’s a difference in the way the orchestrations were done in travel tech and in fintech, because of underlying inherent problem that needs to be solved.
In travel tech, when we were part of MakeMyTrip and Goibibo and we were building things, our responsibility for Kevin was: Kevin is booking hotels in Japan or let’s say in Tokyo, we would be sourcing the price from multiple sources. Kevin can digest 1 second more of latency, but Kevin would be really excited if our deal is $20 better than other platforms. So Kevin, while booking a flight, I don’t think you would mind whether I show you the ticket pricing page in 1 second or 1.5 seconds, right? I don’t think you would mind.
Kevin: I don’t think I would notice, Dhruv. I really don’t.
Dhruv: You would be elated if I’m saving $10 for you.
Kevin: Sure.
Dhruv: Right. So that was our engineering effort in MakeMyTrip and Goibibo. And we tried to ensure we had very smart orchestration mechanisms from multiple sources. And to balance between orchestration and latency, we used a lot of smart caching strategies. And our infra was also designed that way, because we might be latent till 2 seconds, but let’s say Kevin is not seeing flight prices for 35 seconds, Kevin might switch off to some other platform.
Kevin: Especially that. I would notice, Dhruv, that I would notice, yeah.
Dhruv: So it’s a balance between accuracy and how fast we can be. While making products like INDmoney where trading is happening, your responsibility is: usually there is one source, the exchange, where the price change is happening, especially when the day is heavily volatile. So the money that Kevin would put on stake, the equation changes just like this. Within 1 second the price is going up, the price is going down, the market might crash, the market might rise. So in case of INDmoney or similar apps, the responsibility is ensuring the information from source reaches Kevin’s device as fast as possible, as accurately as possible.
So INDmoney engineering was more about utilizing the network aspects in a very granular way, right? Focusing on how much time is my DNS taking to get resolved, what level of compression shall I use, should I use global accelerator or not? And using things like multicast in our infrastructure, because the price is ticking every second. When it ticks, within 50 milliseconds or let’s say within 30, 40 milliseconds, Kevin should be able to see the change and do his trade. Else Kevin might lose money and Kevin would not like my platform. So that has been the difference between the engineering on two sides.
There has been another difference from regulatory point of view. So MakeMyTrip and Goibibo were regulated and we had security certifications and couple of regulations as well. But when you come to fintech, the spectrum of regulation increases. So there is bank regulations, then there is broking industry regulations, there are certificates like ISO. And to cater to that, your scope of FinOps as well as scope of how you run your infrastructure, what kind of monitoring you do, that also increases significantly. But you have to do it because it is part of your business and it is part of your long term business continuity plan.
Kevin: Okay. So then you mentioned compliance earlier and security, and that’s obviously important in all worlds. But when you’re often times there can be friction between moving fast and staying compliant. How do you sort of balance that, and what do you feel is the best way to balance that in a very highly regulated industry like finance?
Chapter 4 — Compliance as a Prerequisite for “Peaceful Sleep”
Dhruv: Yeah. So Kevin, my view here is long term business continuity is always far more important than short term wins, or let’s say short term wins in my budget or short term profits. And long term continuity means trust of two parties, especially trust of the users in my platform and trust of the regulators in my platform.
So in fintech we will always ensure security first, compliance first, and then build our things on top of that. Because if it is a cat and mice game that you are building things which are non compliant, non secure, then in the gap between that cat and mice game, you are highly vulnerable to a lot of things. Losing users’ trust. You are highly vulnerable to different attacks, right? You are highly vulnerable to having a downtime where multiple users may not be able to use your platform at all and there would be user churn. There would be regulatory repercussions and penalties.
So ensure compliance, security and stability first, and then build on top of it. Of course there would be a budgetary penalty because of this. There would be a bloat, but you have to bear that bloat because that is giving you long term continuity.
I’ll take two three examples. So when you are building in fintech, there are requirements like your log has to be retained for these many years, or the level of logging has to be high so that you are logging all the activities that are happening in your infrastructure. And logging as well as the storage is additional cost, right? But I would have to do it because I have to ensure I always have that information during the tenure in which it is required. Another example is how I would do FinOps on this: I might ensure I am storing the information in a better tier of S3. So that’s an optimization I’m doing within the constraint of while being compliant. Another example could be, let’s say, I have to do a lot of identity management and encryption. That is again additional cost, but that additional cost is giving me long term continuity, long term security. And one thing which all of us like is a peaceful sleep, right? So when I’m doing things like this, I can sleep in peace that I am doing a responsible job for my users.
Kevin: Yeah, that’s important. Sleep is important. So I mean, do you use, I’m imagining you use AI for a lot of this stuff. Is it difficult to sort of find the fat in the budget, if you will, without trying to cut the muscle of the security? You talked about how you want to sleep at night, you want everything to be good. Is it difficult to cut that budget but still maintain that security?
Chapter 5 — Identifying “Human Induced Inefficiency” in the Cloud
Dhruv: Yeah, I think so. Our observations have been security is not the primary bloat in your cost. Usually the bloat is human induced inefficiency. So the reason for this is as the platforms are becoming more sophisticated and more user friendly, it’s very easy to spawn a new piece of infrastructure. So in case of INDmoney it was as easy as giving a Slack command that will go to your infrastructure. It will spawn new infra, of course with a human in the loop, just like Kevin, there is a Jira ticket which needs to be approved, and it just spawns. So it can spiral out of control if you are not judicious about it, and that forms a big portion of the inefficiency that you would see.
Security doesn’t. So security costs are predictable. Once you let’s say create a platform for the first time, whether it’s an in house SIEM tool or let’s say you are enforcing encryption on certain databases, those cost rises are not a significant portion of your overall budget. What’s significant bloat is your own inefficiencies, and there FinOps, Cloudgov.ai and other tools, they help big time in monitoring and taking decisions on top of those cost leaks.
So what we do is, there are two parts to it: one is the tech tooling part, and another is creating that culture of FinOps. A lot of companies follow this process called OKR where you do a quarterly plan, 3 month plan. So at INDmoney what we used to do was, there is a total budget of your infrastructure. We would make it more granular at an action level. We used to do something similar in Goibibo and MakeMyTrip as well. So when I say granular level, I mean we would try to find cost per action. So what’s the cost of doing one trade? What’s the cost of booking one hotel or what’s the cost of booking one flight, right? Once you have the target that okay, I’m taking these many cents to get one trade done, what would be the cost 3 months down the line and what would be the cost 6 months down the line? So that evangelization is done between the tech team, the finance team, the product team, and everybody has a handshake on it, and then that’s how the things proceed. So that’s the human side of it.
How FinOps tooling side helps is: if your infrastructure is really lean and let’s say you only have one or two accounts, you can humanly manage those cost with a lean team, right? But in fintech, even in travel tech, you usually have multiple accounts. You usually have multiple businesses each running in isolation like a specialized business unit, and that kind of magnifies the complexity by 5x, 10x based on how complex your infrastructure is.
Chapter 6 — Moving from Manual Scripts to Agentic AI Scalability
Dhruv: That’s where Agentic AI comes into play, not as a replacement of the human but something supplementing the power of the human. So they are able to derive decisions by observing all of your infra. They’re able to take decisions as well. But I think the point where they are able to take all the decisions is not there yet. But they can take partial decisions on infrastructure which is of less severity, and they can guide the human on the infrastructure where the cost of taking decision is high.
Kevin: Yeah. Was there a moment when you sort of just went, “Oh, wow. An AI agent could handle a complex FinOps task better than a manual human script.”
Dhruv: Yep. So I would not say better. It was expected. A good thing about a program or an agent is scalability. I would say better if the outcomes were such that a human could not have been able to do it. That was not the case. A human given due time would be able to do the same things, but the wow factor is in how fast and how much. That’s where the Agentic comes into play.
So our wow moment was replicating the same FinOps goodness across multiple accounts with diminished effort. So your Agentic tooling, FinOps tooling, is running into one account, integrating another account or let’s say a parallel infra, was 1 by 10th of the effort. For a human it would have been a parallel effort, or at least they would have to create the tooling to reduce the effort. But when you are running Agentic systems, the effort is incrementally lower and lower, and the output or the outcome that you are getting is significantly higher. So that’s the wow factor that I noticed.
Kevin: Okay.
Dhruv: And that’s the exact reason for which a FinOps tool starts making budgetary sense as well. Because if a $1,000 tool is saving me only $1,000 of cost, then I’m at a net zero saving, right? There is no sense of me adopting that product, rather it’s just adding complexity to my stack. But when a $1,000 tool is able to save $1,000 for me 20 times, then the ROI that I am receiving from that tooling is really important and very good for my long term budgeting. So that’s where they start making sense.
Kevin: Do you think we’ll reach a point of self healing infrastructure that kind of manages its own costs?
Chapter 7 — Self Healing Infra: Sandbox vs. Production Reality
Dhruv: Can’t say right now, but of course the way AI is evolving and the problems are being fed as a loop to the systems, there would be few inflection points where things would start working that way. Right now I can say we can use it for a fully AI driven self healing system in not production critical systems. Let’s say I have pre production and developer environments where I might be leaking cost by running them all the time and running unoptimized instances. This even right now is a use case for fully self healing infra running, and fully cost optimized infra running fully through AI agents.
But would I be willing to do that on a production system where one lakh people would be coming and doing trades in 1 hour from now? I think there I would still want to have a human in the loop for the responsibility part and for the compliance part. So the very sensitive infrastructure, the AI tooling might be giving the direction and the decisioning, the human in the loop would still be approving the actions, right? Eventually it might become fully autonomous, but right now for certain critical systems people would still want a human in the loop.
Kevin: Yeah, I would imagine especially for doing certain things like, like you said, being in the space for INDmoney and being in these spaces where compliance and regulatory concerns are high, to take a human out of the loop completely, boy, you’d have to have a lot of confidence, right?
Dhruv: Yep.
Kevin: And I would imagine we are fairly far from that even though we’re moving fast.
Dhruv: I would not say far or close, because AI keeps on surprising us every month. But true, it’s not just a technical problem here. It’s an amalgamation of what’s technically feasible and what compliance and regulations say. So you would have to fit in the technical solution within the framework of compliance, and then arrive at the most optimal solution. Had it been some other industry, I might have been more adventurous, and you know, as we talked about fully autonomous AI systems, those experiments could be more lucrative to do there as compared to fintech. Being very honest about this.
Kevin: Yeah, that makes sense to me. I could see that too. Yeah, if you were doing something that was a little less intense in that respect. A lot of CTOs struggle with rising cloud bills, complex tech stacks. Is there anything that you would recommend a CTO do now to sort of prepare for this Agentic shift if they haven’t started preparing yet?
Chapter 8 — Advice for CTOs: Starting the FinOps Journey
Dhruv: So obviously first is a zero or one situation. If you do not have a FinOps team, let’s not even talk about the tooling. If you don’t have a set of people who are actively looking at cost, that’s the first very basic step. And that basic step with a basic team and a basic FinOps tooling would ensure you are removing the obvious leakages from your infrastructure. You will find unutilized infrastructure. You’ll find the scaling policies are really optimistic. You are wasting the infrastructure. So that would save initial cost and reduce.
Then based on the complexity of the stack, people can do a small PoC on a small portion. Because once you get encouraged by seeing result on small portion of your infrastructure, then it is easy to convince a lot of people. Because while adopting a FinOps tooling, there would be budget as well, right? And you would justify budget by saying that this is going to save us 5x cost. I would be spending $100 extra now, but this $100 is going to save $1,000 for me in near future. When you do a proof of concept which proves this on one part of the infrastructure, then it’s really easy to replicate on other part of your infrastructure.
And secondly, AI replaces humans, right? This is being talked about. So that’s the reality. But AI also supplements humans. So by observing what AI agents are able to do in your FinOps and cloud optimization, you will be able to better forecast what kind of infrastructure management and FinOps human team you would want to create. Because if you are not using any tooling, the way you would create your team versus while you are using tooling and then you are creating the team, the second team would be leaner in 100% of the cases. So that helps you forecast your hiring requirements as well.
Kevin: Yeah, that’s a good point. There’s obviously a lot of concern with AI taking jobs, but you’re basically kind of saying that there’s a place for the human in the loop still. And for now anyway, there is a place for the human in the loop. Do people in the FinOps space, do people in the fintech space, do they need to really make sure they understand how these AI systems operate in order to stay relevant, on top of things, to not be replaced?
Dhruv: If we talk about fintech, AI is not going to replace these humans, especially in regulated spaces. As I talked, AI is going to magnify what they can do. And yes, you’re right, they need to be on top of what model to utilize, right? What could be the security implications, because all of the AI agents, just like my normal code is susceptible to some form of attack, my Agentic AI is also susceptible to some form of attack and some form of malicious behavior, right? So what are my guardrails in that case? How am I enforcing guardrails? One simple example of guardrails is, let’s say I’m using Agentic AI to take decisions. I may allow it to take downscaling decisions, but I may not allow it to take decisions where it is shutting down certain part of infra, or let’s say I may not allow decisions related to databases. So there is still a role of a human in the loop who has the business context, who understands what could be the predictions for the next quarter, what can increase, what can decrease, what can happen if market suddenly spikes due to XYZ reasons. So it’s a handshake. It’s not a replacement, in my opinion.
Kevin: Yeah, I agree in a lot of cases. I mean, obviously there’s some cases where AI can replace people, but as a general rule I kind of always look at it as a tool. It’s a tool to make us better and more efficient. I think that’s the thing I like about the AI tools that I use is just the sheer efficiency that I work with now.
Dhruv: Yeah, I think a very good use case is your podcast, right? So right now I’m not talking to an AI avatar of Kevin, but I’m sure Kevin as a human is utilizing a lot of AI tools in his storytelling or post processing and pre processing, right?
Kevin: Yep, absolutely. And you’re right, it’s a good example. It’s not replacing me, but it’s making me way more efficient and able to do more, and less pulling out of my hair, which is nice. That’s a good outcome.
Dhruv: That is a good outcome.
Kevin: I don’t have a lot left to pull out. So you’ve recently stepped aside from INDmoney. What’s next for Dhruv? What’s on the horizon? Anything you can tell us about?
Chapter 9 — What’s Next for Dhruv Pathak?
Dhruv: Kevin, I’m right now slowing down and observing the ecosystem. So good part is right now we are in a cusp and an inflection point where the AI ecosystem, the content ecosystem, travel, finance, all of the ecosystems are so vibrant and a lot of things are happening around. So I’m observing all the exciting parts and figuring out what I would pursue next. Right now I’m advising few startups on their initial journeys and that’s an equally exciting thing to do.
Kevin: Absolutely. Is it kind of nice to take a little breather?
Dhruv: Yeah, it is. And plus, I’m able to touch multiple different types of products. So that’s also interesting.
Kevin: Yeah, that’s exciting.
Dhruv: And who knows, one of the outcomes could be I could be a podcast host like you in future.
Kevin: You never know. You’re a natural. I’ll say that. Dhruv, you’re a natural.
Dhruv: Thanks.
Kevin: All right. Well, Dhruv Pathak, thanks so much for joining us today. I really appreciate all the insight, and yeah, I’d love to chat with you again sometime.
Dhruv: Sure, Kevin. Be in touch.
Kevin: And thank you all for joining us today. If today’s conversation got you thinking about how to make your FinOps truly self driving, check out Cloudgov.ai, where Agentic AI turns cloud cost management into autonomous action. Especially if you’re tired of legacy vendors that overpromise and underdeliver, it’s time to move to a platform that actually evolves with you. Visit Cloudgov.ai and experience how consistent innovation and AI powered automation are redefining the future of FinOps. Until next time, keep everything well operated.


