Cloud risk was one of the top-level enterprise risks for the company, and we were not alone in identifying it as the top-level risk. Here is a story that will bring this closer to home — a true story.
Alex I cannot believe that they would do that.
Sasha Do what?
Alex Create a link from dev to prod.
Sasha Oh! Why would they do that?
Alex Something about the commercial license middleware issue. We didn’t buy an additional license for the dev environment, so they had to test in production.
Everyone smiled; a different expression wasn’t appropriate — everyone knew what was coming. I led that response, stepping in for my lead, who was on vacation in Hawaii; he deserved the time off. We worked through the problem with all hands on deck, locked in a tiny war room: glass covered with sheets of paper, a whiteboard with timelines drawn in different colors, threat intelligence on attacker tactics, and the known facts at that point. You would understand if you have seen one — no one left the room. People walk by, curious about what is happening, but they are not supposed to look. A disclosure list was on the whiteboard. You don’t want to be on one; if you can avoid it, it feels special, but the novelty evaporates quickly if it ever goes wrong, and it does. I was unfortunate: I was always there on every disclosure list.
Incident response teaches you to remain calm; you have to keep your head level. After the initial short exchange there was no blame; the focus shifted to what was at hand. An intruder had used a zero-day to exploit an unprotected development environment. They had a shell beaconing home, and they were now sitting on the production systems with critical data — the intruders didn’t know that yet. The following sixteen hours were challenging, a chess game with a timer. We eventually evicted the intruders and secured the system without any data exposure. It is quite a story, but it is a very different conversation. It was also one of the most rewarding sixteen hours of my life.
The incident postmortem revealed the development team had created a temporary link between the development and production cloud environments. Only a handful of people knew about it, and it wasn’t documented. This is not an isolated incident or practice; it happens across most organizations for one reason or another. The focus is on value generation, and the faster we go, the quicker the business generates value. We didn’t know, and it hurt us — almost did us in, in this case.
So what is the problem?
Lack of holistic lifecycle visibility. The dynamic nature of the cloud, the speed of development, and the proclivity for creating fast-moving, highly autonomous teams more often than not lead to this situation. Organizations of any size will, over time, lose track of what is in the cloud and of resource relationships and interdependencies. “No one knows what is in the cloud” holds quite true. Cloud risk grows; we simply cannot protect what we do not know. The end result is that remediation is quite literally broken and cloud risk goes unmanaged.
Most cloud environments have multitudes of these configurations, created just in time to accomplish a critical task and often forgotten. Every such configuration adds to the cloud risk: a hidden, undocumented, unapproved access. Silos between functional teams, tools, and capabilities add to the confusion. As a result, misconfigurations are missed, and vulnerabilities are not prioritized across the lifecycle of a capability. Gartner’s reporting notes that cloud misconfiguration, along with application vulnerabilities, are the two biggest sources of incidents. Everyone feels the pain, is overworked, and cannot seem to handle all the data generated by a plethora of tools. The tendency is to add more, generating more data for the teams to consume, and every tool adds workload to already oversubscribed teams. With no end in sight, we return to add yet another tool that might solve the problem, and find ourselves with a YAP — yet another tool problem.
As we add more tools, remediation is prioritized along functional silos, and teams are not working on the highest-priority risk. This adds more work, makes teams inefficient, and often adds to the overall cloud risk.
An alternative: the lifecycle view
We propose an alternative, starting with full visibility of the cloud — the lifecycle view. Get the basics right; get the visibility right. Data extracted from all environments — dev, stage, prod. All functional areas: design, development, and production; code, applications, and cloud infrastructure. Every aspect that is needed to secure the cloud is tied together through relationships and connections — network access, policy-based serverless and APIs, running applications, code and infrastructure — in one lifecycle view, and all other data is an overlay on that basic visibility. Remediation is prioritized horizontally, with the highest risk getting the priority, but also across the full cloud portfolio, so teams are always working on the highest-risk areas. We can see a vision where we might not want to add more tools. We can get this right.
Start building a true risk picture; with that, we can handle cloud risk. The lifecycle view is powerful because it starts breaking the functional silos, and teams can work more efficiently and more quickly. We believe this has the potential to make risk not just a word we throw around but one we use to make the appropriate business decisions — because now we have the basics right, and we know what is in our environment, and not just in production.