AI explainability is easy to underestimate when AI is used for small, everyday tasks. If a model recommends the wrong article, produces a weak summary, or gives a slightly strange answer, the consequences are usually limited. We may be annoyed, but we can move on.
The problem begins when AI becomes part of decisions that people cannot simply ignore. A loan application, an insurance claim, a medical recommendation, a hiring process, a court case, or an autonomous system failure all create a very different expectation. In those situations, people need more than an output. They need a way to understand what influenced it, whether it was fair, and whether it can be challenged.
That is why AI explainability should not be treated as decoration around a model. A nice paragraph next to a prediction may improve the interface, but it does not automatically create accountability. The explanation has to match the decision, the risk, and the person who needs to use it.
As regulation catches up with AI adoption, companies will have to think about explainability much earlier in the product lifecycle. Not after deployment, not only when lawyers ask for it, and not as a checkbox. It has to be part of how AI systems are designed, tested, documented, and governed.
https://mkdev.me/posts/explaining-ai-explainability-vision-reality-and-regulation
The problem begins when AI becomes part of decisions that people cannot simply ignore. A loan application, an insurance claim, a medical recommendation, a hiring process, a court case, or an autonomous system failure all create a very different expectation. In those situations, people need more than an output. They need a way to understand what influenced it, whether it was fair, and whether it can be challenged.
That is why AI explainability should not be treated as decoration around a model. A nice paragraph next to a prediction may improve the interface, but it does not automatically create accountability. The explanation has to match the decision, the risk, and the person who needs to use it.
As regulation catches up with AI adoption, companies will have to think about explainability much earlier in the product lifecycle. Not after deployment, not only when lawyers ask for it, and not as a checkbox. It has to be part of how AI systems are designed, tested, documented, and governed.
https://mkdev.me/posts/explaining-ai-explainability-vision-reality-and-regulation
mkdev.me
Decoding AI Explainability: Vision, Reality & Regulation
AI can feel magical, but when decisions affect health, justice, or safety, we need more than magic—we need explanations. Paul Larsen breaks down what “explainable AI” really means, why different stakeholders need different kinds of “why,” and how this series…
Back in 2023, we tested AWS App Runner as a simpler way to deploy containers without managing all the usual infrastructure around ECS.
In 2026, App Runner is closed to new customers and AWS recommends ECS Express Mode instead. Watch the video to see where the idea worked — and where it didn’t: https://www.youtube.com/watch?v=E6E6HtrLs98
In 2026, App Runner is closed to new customers and AWS recommends ECS Express Mode instead. Watch the video to see where the idea worked — and where it didn’t: https://www.youtube.com/watch?v=E6E6HtrLs98
YouTube
Is AWS AppRunner the worst way to run containers?
AWS AppRunner is one of the latest additions to a billion ways to run containers on AWS. Is it any good? Let's find out!
DevOps Accepts Episode about "voice-to-gpt" project: https://mkdev.me/podcast
Pablo's video about this tool https://www.youtube.com/…
DevOps Accepts Episode about "voice-to-gpt" project: https://mkdev.me/podcast
Pablo's video about this tool https://www.youtube.com/…
If you’re one of our Spanish-speaking subscribers and want to dive into containers, we have a free video course for you!
It covers Docker, docker-compose, Docker Swarm, Podman, Buildah, Firecracker and more — all in Spanish and completely free: https://www.youtube.com/playlist?list=PLNXwhzx0-DmRlCz9lKPLRWdv5visGlNNT
It covers Docker, docker-compose, Docker Swarm, Podman, Buildah, Firecracker and more — all in Spanish and completely free: https://www.youtube.com/playlist?list=PLNXwhzx0-DmRlCz9lKPLRWdv5visGlNNT
🔥1
Cloud lock-in isn’t only about proprietary APIs and managed services.
Economics can create lock-in too.
Historically, one of the clearest examples was data egress: getting data into a cloud was often cheap or free, while getting large amounts of it back out could become expensive.
In 2026, the picture is changing. AWS, Google Cloud and Azure all offer programs that waive eligible egress charges when customers completely migrate away. In the EU, the Data Act has already started changing the rules around cloud switching, with switching charges due to be fully prohibited from January 2027.
But there is an important distinction: making it cheaper to leave a cloud does not mean data transfer itself has become free.
Applications still generate network costs between zones, regions, services and the public internet. Those costs can influence architecture just as much as compute or storage pricing.
The cloud is becoming easier to leave. Understanding the cost of moving data while you are still there remains just as important.
https://mkdev.me/posts/the-biggest-cloud-scam
Economics can create lock-in too.
Historically, one of the clearest examples was data egress: getting data into a cloud was often cheap or free, while getting large amounts of it back out could become expensive.
In 2026, the picture is changing. AWS, Google Cloud and Azure all offer programs that waive eligible egress charges when customers completely migrate away. In the EU, the Data Act has already started changing the rules around cloud switching, with switching charges due to be fully prohibited from January 2027.
But there is an important distinction: making it cheaper to leave a cloud does not mean data transfer itself has become free.
Applications still generate network costs between zones, regions, services and the public internet. Those costs can influence architecture just as much as compute or storage pricing.
The cloud is becoming easier to leave. Understanding the cost of moving data while you are still there remains just as important.
https://mkdev.me/posts/the-biggest-cloud-scam
mkdev.me
Cloud Scam Exposed: Uncover Hidden Egress Fees | mkdev
Every time you communicate with a machine in different availability zones, you have to pay. Every time there is an egress, you have to pay. Every time you exit a cloud, you have to pay. You always end up paying. This is the biggest scam in the history of…
Misconfigured RBAC, weak network policies or poorly managed secrets can leave a Kubernetes cluster exposed.
Our In-Depth Kubernetes Security Audit helps uncover these risks and gives your team a practical path to address them. Explore the audit and talk to us about your setup: https://mkdev.me/b/audits/kubernetes-security-audit
Our In-Depth Kubernetes Security Audit helps uncover these risks and gives your team a practical path to address them. Explore the audit and talk to us about your setup: https://mkdev.me/b/audits/kubernetes-security-audit
mkdev.me
Kubernetes Security Audit | mkdev audits for business
Navigating the web of Kubernetes security demands a nuanced understanding and a meticulous eye for detail. That's where our expert team comes into play.
There are really two different problems hiding behind the term “AI explainability.”
The first is understanding how a model behaves in general. Global explainability methods can tell us which features tend to matter across a population and are particularly useful for developers who want to understand or debug a model.
The second is explaining one particular decision. Why was this loan application rejected? Why did this model produce this prediction? Local explainability methods such as LIME and SHAP try to answer those questions by building simpler approximations around individual cases.
The distinction matters because a population-level explanation doesn't necessarily tell you why something happened to one person. And a local approximation, however useful, isn't the same thing as opening up the original black box.
Explainability therefore isn't one technology solving one problem. It's a collection of approaches with different strengths, limitations, audiences and purposes.
We explored these questions in our article on AI explainability, and the distinction remains an important one for businesses working with increasingly complex AI systems: https://mkdev.me/posts/explaining-ai-explainability-the-current-reality-for-businesses
The first is understanding how a model behaves in general. Global explainability methods can tell us which features tend to matter across a population and are particularly useful for developers who want to understand or debug a model.
The second is explaining one particular decision. Why was this loan application rejected? Why did this model produce this prediction? Local explainability methods such as LIME and SHAP try to answer those questions by building simpler approximations around individual cases.
The distinction matters because a population-level explanation doesn't necessarily tell you why something happened to one person. And a local approximation, however useful, isn't the same thing as opening up the original black box.
Explainability therefore isn't one technology solving one problem. It's a collection of approaches with different strengths, limitations, audiences and purposes.
We explored these questions in our article on AI explainability, and the distinction remains an important one for businesses working with increasingly complex AI systems: https://mkdev.me/posts/explaining-ai-explainability-the-current-reality-for-businesses
mkdev.me
AI Explainability: Complexity, Trust & Business Impact
In the second article of his explainable AI series, Paul Larsen looks at what today’s XAI tools really deliver for different stakeholders—from users to regulators—and where they still fall short for trust, liability and high-risk decisions.
If Linux networking still feels like a collection of mysterious interfaces and commands, this one is worth revisiting.
Learn how teaming, Linux Bridge, tap interfaces and Traffic Control work together for fault tolerance and bandwidth management.
Read more: https://mkdev.me/posts/how-networks-work-part-two-teaming-for-fault-tolerance-bandwidth-management-with-traffic-control-tap-interfaces-and-linux-bridge
Learn how teaming, Linux Bridge, tap interfaces and Traffic Control work together for fault tolerance and bandwidth management.
Read more: https://mkdev.me/posts/how-networks-work-part-two-teaming-for-fault-tolerance-bandwidth-management-with-traffic-control-tap-interfaces-and-linux-bridge
mkdev.me
Linux Teaming: Fault Tolerance & Traffic Control | mkdev
We’re going to talk about how Linux Bridge, tap interfaces and Linux Traffic Control work and what you need them for as well as virtualization using these tools.
In the 97th mkdev dispatch Kirill explains the role Terraform has in this new age of AI agents. Also inside: AWS fast networking, Aurora DSQL pricing and more!
https://mkdev.me/posts/terraform-in-the-ai-agents-age-97
https://mkdev.me/posts/terraform-in-the-ai-agents-age-97
mkdev.me
Terraform: Why It Still Matters in the AI Era | mkdev
In the 97th mkdev dispatch Kirill explains the role Terraform has in this new age of AI agents. Also inside: AWS fast networking, Aurora DSQL pricing and more!
Imagine Google Cloud tells you you're spending roughly €500 a year on something you don't remember creating.
The obvious next step is to open FinOps Hub. You can inspect recommendations, look for potential savings and check where the spending is coming from.
But then you discover that the cost isn't an application server at all. It's infrastructure created to provide VPC connectivity for Cloud Run.
That's where the interesting part of FinOps starts.
Today, Direct VPC egress is Google's recommended approach for many Cloud Run workloads and avoids the compute cost of running Serverless VPC Access connector instances. It's a small architectural change that can remove an entire category of unnecessary spending.
We walk through this example, along with FinOps Hub, CUDs, cost allocation and billing analysis, in our Google Cloud FinOps article.
https://mkdev.me/posts/gcp-finops-hub-the-key-to-mastering-your-finances-on-google-cloud
The obvious next step is to open FinOps Hub. You can inspect recommendations, look for potential savings and check where the spending is coming from.
But then you discover that the cost isn't an application server at all. It's infrastructure created to provide VPC connectivity for Cloud Run.
That's where the interesting part of FinOps starts.
Today, Direct VPC egress is Google's recommended approach for many Cloud Run workloads and avoids the compute cost of running Serverless VPC Access connector instances. It's a small architectural change that can remove an entire category of unnecessary spending.
We walk through this example, along with FinOps Hub, CUDs, cost allocation and billing analysis, in our Google Cloud FinOps article.
https://mkdev.me/posts/gcp-finops-hub-the-key-to-mastering-your-finances-on-google-cloud
mkdev.me
Optimize Google Cloud Costs with GCP FinOps Hub | mkdev
Managing your Google Cloud costs is crucial for any business and finances. In this tutorial, we’ll cover everything how to use GCP FinOps, cost management and how to setup Direct VPC egress to reduce Cloud Run costs. Learn how to optimize your spending and…
CI/CD, infrastructure as code, Kubernetes and observability are powerful building blocks. Platform Engineering is about turning them into a coherent experience that helps teams ship efficiently at scale.
Read about our approach to Platform Engineering and arrange a call to discuss your setup: https://mkdev.me/b/consulting/platform-engineering
Read about our approach to Platform Engineering and arrange a call to discuss your setup: https://mkdev.me/b/consulting/platform-engineering
mkdev.me
Platform Engineering Consulting | mkdev
Schedule a call to receive the Platorm Engineering consultation
The interesting part about modern image models might not be image quality anymore.
It’s iteration.
Generating one impressive image is easy. Generating 50 or 100 versions, remembering what you disliked about previous attempts, making targeted corrections and gradually converging on something useful is a different problem.
That’s why combining Claude Code with image-generation models turned out to be more interesting than simply using another image-generation UI.
Claude can maintain the context of the task and use Nano Banana Pro or GPT Image 2 as tools. It can look at the output, notice that an object is positioned strangely or that the result doesn’t quite satisfy the request, and take another shot.
That turns image generation from a sequence of isolated prompts into something closer to an iterative creative workflow.
We explored the approach in this article and made the skills public:
https://mkdev.me/posts/unlimited-image-generation-with-nano-banana-pro-gpt-image-2-and-claude-code-skills
It’s iteration.
Generating one impressive image is easy. Generating 50 or 100 versions, remembering what you disliked about previous attempts, making targeted corrections and gradually converging on something useful is a different problem.
That’s why combining Claude Code with image-generation models turned out to be more interesting than simply using another image-generation UI.
Claude can maintain the context of the task and use Nano Banana Pro or GPT Image 2 as tools. It can look at the output, notice that an object is positioned strangely or that the result doesn’t quite satisfy the request, and take another shot.
That turns image generation from a sequence of isolated prompts into something closer to an iterative creative workflow.
We explored the approach in this article and made the skills public:
https://mkdev.me/posts/unlimited-image-generation-with-nano-banana-pro-gpt-image-2-and-claude-code-skills
mkdev.me
Unlimited Image Generation: Nano Banana Pro & GPT Image 2
Nano Banana Pro and OpenAI's GPT Image 2 are top-tier image gen models right now — and Kirill wired both into Claude Skills. 100+ icon iterations, 4K control, self-critiquing generations, and sane context handling. $45 well spent.
A useful Cloud Run distinction:
Services → request-driven applications
Jobs → run-to-completion workloads
And when the workload can be divided, Cloud Run Jobs can execute multiple tasks in parallel.
We break down the idea with a practical example here: https://www.youtube.com/watch?v=n8GyTp-kP_M
Services → request-driven applications
Jobs → run-to-completion workloads
And when the workload can be divided, Cloud Run Jobs can execute multiple tasks in parallel.
We break down the idea with a practical example here: https://www.youtube.com/watch?v=n8GyTp-kP_M
YouTube
How to use Google Cloud Run Jobs for background tasks
In this video we are going to learn Cloud Run Jobs, release as GA this week and how now everything is different
AWS Load Balancer Controller with EKS, our free webinar: http://mkdev.me/webinars/aws-lb-controller-101
Check out mkdev dispatch, a bi-weekly…
AWS Load Balancer Controller with EKS, our free webinar: http://mkdev.me/webinars/aws-lb-controller-101
Check out mkdev dispatch, a bi-weekly…
👍1
Want to understand containers beyond Docker? Our free Dockerless course takes you through OCI, open container standards, and the fundamentals that make modern containers work.
Article series: https://mkdev.me/posts/what-s-wrong-with-docker-introduction-to-the-dockerless-course
Video: https://www.youtube.com/playlist?list=PLozcbFx8FoPH30kYPbPuPsvxASWoLo9XB
Article series: https://mkdev.me/posts/what-s-wrong-with-docker-introduction-to-the-dockerless-course
Video: https://www.youtube.com/playlist?list=PLozcbFx8FoPH30kYPbPuPsvxASWoLo9XB
mkdev.me
Dockerless Course: Rethink Containers & Open Standards
We’ve used Docker for so long, we’ve started calling everything by its name—even when it’s not Docker. In the introduction to our most popular course, Kirill Shirinkin invites you to rethink what containers really are, why the industry is moving beyond Docker…
🔥1🎉1
For a while, microservices felt like the inevitable destination of every serious application.
Then teams discovered the other side of the equation: more services also mean more network calls, more APIs, more deployments, more dependencies, and more opportunities for things to fail in ways that are difficult to reproduce locally.
Google’s Service Weaver experimented with an interesting alternative. Developers could structure a Go application as a set of components without immediately committing every component to a separate service. Deployment topology could be decided later.
Service Weaver didn’t become the future of application development, and the project has since been archived. But the problem it was trying to solve hasn’t disappeared.
The interesting part of Service Weaver in 2026 isn’t the framework itself. It’s the question it leaves behind: should our code architecture really be so tightly coupled to our deployment architecture?
https://mkdev.me/posts/service-weaver-monolithic-or-microservice
Then teams discovered the other side of the equation: more services also mean more network calls, more APIs, more deployments, more dependencies, and more opportunities for things to fail in ways that are difficult to reproduce locally.
Google’s Service Weaver experimented with an interesting alternative. Developers could structure a Go application as a set of components without immediately committing every component to a separate service. Deployment topology could be decided later.
Service Weaver didn’t become the future of application development, and the project has since been archived. But the problem it was trying to solve hasn’t disappeared.
The interesting part of Service Weaver in 2026 isn’t the framework itself. It’s the question it leaves behind: should our code architecture really be so tightly coupled to our deployment architecture?
https://mkdev.me/posts/service-weaver-monolithic-or-microservice
mkdev.me
Google's Service Weaver: Unify Monoliths & Microservices
Pablo Inigo Sanchez examines the complexities of microservices deployment and highlights Google's Service Weaver, a tool designed to streamline the development of distributed applications, addressing many common challenges.
Using Google Cloud doesn't automatically mean you're using the right services in the right way.
Our GCP Audit helps uncover unnecessary costs, security gaps, reliability issues and opportunities to simplify your infrastructure.
Take a look at what we cover, and schedule a conversation with our team: https://mkdev.me/b/audits/google-cloud-platform
Our GCP Audit helps uncover unnecessary costs, security gaps, reliability issues and opportunities to simplify your infrastructure.
Take a look at what we cover, and schedule a conversation with our team: https://mkdev.me/b/audits/google-cloud-platform
mkdev.me
Google Cloud Platform | mkdev audits for business
As part of Google Cloud Platform audit and assessment, we take a deep review of your setup from security and high availability to cost and automation. We help you to decide what component to use in every case for your business.
There are two GenAI security concepts that are easy to mix up: jailbreaks and prompt injection.
A jailbreak generally tries to make an AI system ignore its restrictions and produce something it shouldn’t.
Prompt injection can go further. The goal may be to manipulate an AI system into taking actions or accessing systems and data that the attacker should never be able to reach.
As GenAI gets connected to more tools, APIs, and business systems, that distinction becomes increasingly important. The potential consequences move from “the chatbot said something bad” to data leakage, compromised systems, and operational disruption.
Our guide explains these risks from a product manager’s perspective, alongside another major issue: how employees and users handle sensitive data with GenAI tools.
Read the full article on mkdev: https://mkdev.me/posts/genai-security-risks-for-product-managers-dd73bdc2-4f2e-4227-93b3-375da081d906
A jailbreak generally tries to make an AI system ignore its restrictions and produce something it shouldn’t.
Prompt injection can go further. The goal may be to manipulate an AI system into taking actions or accessing systems and data that the attacker should never be able to reach.
As GenAI gets connected to more tools, APIs, and business systems, that distinction becomes increasingly important. The potential consequences move from “the chatbot said something bad” to data leakage, compromised systems, and operational disruption.
Our guide explains these risks from a product manager’s perspective, alongside another major issue: how employees and users handle sensitive data with GenAI tools.
Read the full article on mkdev: https://mkdev.me/posts/genai-security-risks-for-product-managers-dd73bdc2-4f2e-4227-93b3-375da081d906
mkdev.me
GenAI Security Risks for Product Managers
The third article of this series by Paul Larsen warns product managers about the major cybersecurity risks of GenAI—like data leaks, prompt jailbreaks, and injection attacks—and offers practical steps to keep AI use productive without endangering company…
Kirill Shirinkin has spent decades across software development, infrastructure and DevOps. For nearly a decade, his own development setup barely changed.
Then AI arrived.
Now he shares what he’s learned from rebuilding his workflow around AI: how he works with coding agents, manages context, makes architecture decisions, organizes parallel work, automates reviews and deployments, and decides what should still remain firmly in the engineer’s hands. He also looks at the less glamorous side of working at AI speed: avoiding bad technical decisions, keeping projects under control, and staying sane when the amount of work you could do suddenly feels almost limitless.
https://mkdev.me/posts/the-agentic-engineering-myth-1-year-of-coding-with-ai
Then AI arrived.
Now he shares what he’s learned from rebuilding his workflow around AI: how he works with coding agents, manages context, makes architecture decisions, organizes parallel work, automates reviews and deployments, and decides what should still remain firmly in the engineer’s hands. He also looks at the less glamorous side of working at AI speed: avoiding bad technical decisions, keeping projects under control, and staying sane when the amount of work you could do suddenly feels almost limitless.
https://mkdev.me/posts/the-agentic-engineering-myth-1-year-of-coding-with-ai
mkdev.me
The Agentic Engineering Myth - 1 Year of Coding with AI
A practical guide to agentic engineering in 2026: AI coding agents, tools, workflows, context management, architecture, automation, and shipping software faster.
🔥1
In the 98th mkdev dispatch, Pablo talks about how AI is creating a new digital divide where access to powerful tools depends not on infrastructure or skills, but on geography, provider restrictions, and permission to participate in the AI economy. Also inside: scaling Terraform across many Teams and more!
https://mkdev.me/posts/the-ai-divide-is-no-longer-about-access-to-the-internet-98
https://mkdev.me/posts/the-ai-divide-is-no-longer-about-access-to-the-internet-98
mkdev.me
AI Divide: When Access Becomes a Privilege | mkdev
In the 98th mkdev dispatch, Pablo talks about how AI is creating a new digital divide where access to powerful tools depends not on infrastructure or skills, but on geography, provider restrictions, and permission to participate in the AI economy. Also inside:…
AWS gives you plenty of ways to analyze cloud spend, but the analysis is only as useful as the metadata behind it.
For smaller AWS setups, you don’t necessarily need an elaborate tagging framework. A simple baseline of environment, workload, and name can already make it much easier to understand which application or environment is responsible for a particular part of the bill.
Terraform’s default_tags can help apply that baseline consistently across supported resources. The other important step is easy to miss: tags need to be activated as Cost Allocation Tags before you can properly use them for cost analysis.
A relatively small amount of tagging discipline can make Cost Explorer considerably more useful.
Our article walks through a simple approach to getting started with AWS cost allocation tags: https://mkdev.me/posts/control-aws-costs-with-these-3-cost-allocation-tags
For smaller AWS setups, you don’t necessarily need an elaborate tagging framework. A simple baseline of environment, workload, and name can already make it much easier to understand which application or environment is responsible for a particular part of the bill.
Terraform’s default_tags can help apply that baseline consistently across supported resources. The other important step is easy to miss: tags need to be activated as Cost Allocation Tags before you can properly use them for cost analysis.
A relatively small amount of tagging discipline can make Cost Explorer considerably more useful.
Our article walks through a simple approach to getting started with AWS cost allocation tags: https://mkdev.me/posts/control-aws-costs-with-these-3-cost-allocation-tags
mkdev.me
Optimize AWS Billing: 3 Key Allocation Tags | mkdev
Learn how to optimize AWS costs with a comprehensive tagging strategy using Terraform. Properly tagging your AWS resources can enhance cost analysis in AWS Cost Explorer, providing insights for cloud cost monitoring and management. Discover simple techniques…
👍1
We’ve been using Infrastructure as Code for over a decade. The tools have changed, but the principle hasn’t: infrastructure should be understandable, version controlled, tested and automated.
Learn how we work and start a conversation: https://mkdev.me/b/consulting/iac
Learn how we work and start a conversation: https://mkdev.me/b/consulting/iac
mkdev.me
Infrastructure as Code & GitOps consultation for business | mkdev
Schedule a call to receive the Infrastructure Deployment consultation for Advanced level developers from industry experts
🎉1