Every data profesional hits a point in their career, where numbers start feeling a little “too confident”. You look at a single prediction or coefficient and instinctively know there’s more uncertainty hiding underneath it. I remember reaching that stage clearly, when dashboards looked clean but decisions still felt risky. The data was speaking, but it wasn’t telling the full story. That’s usually when curiosity kicks in and you start asking deeper questions about confidence, not just accuracy. Bayesian modeling enters the picture right there, not as a replacement for what you know, but as an evolution of it. It gives uncertainty a voice instead of sweeping it under the rug. And once you hear that voice, it’s hard to ignore it going forward.
What you will learn: In this edition, we’re exploring Bayesian modeling and how to think about uncertainty in a more realistic, practical way using PyMC3. By the time you’re done, you’ll have a clear intuition for what Bayesian thinking really means, why it’s so useful in day-to-day data work, and how it changes the way you interpret results. We’ll also explore how PyMC3 supports this mindset in a structured but approachable way.
Read Time: 9 minutes
Source: Sahir Maharaj (https://sahirmaharaj.com)
The first thing Bayesian modeling asks of you has nothing to do with code or tools. It asks you to rethink how you define “an answer.” Most of us were trained to chase a single value: the average, the coefficient, the forecast. That approach works, until it doesn’t. Bayesian thinking gently pushes back and asks a different question: “What range of outcomes actually makes sense here?” That shift sounds subtle, but it fundamentally changes how you reason about data. In traditional approaches, uncertainty is usually something you acknowledge briefly and then move past. You might calculate a confidence interval, mention it in passing, and still focus most of the conversation on the point estimate. In Bayesian modeling, uncertainty is not a footnote. It’s the main story. Every parameter has a distribution, and that distribution carries meaning about how sure (or unsure) you really are.
What I’ve found interesting is that this mirrors how people already think, even if they don’t use statistical language. Stakeholders naturally ask questions like “How confident are we?” or “What’s the risk if this goes wrong?” Bayesian models simply give you a structured way to answer those questions instead of hand-waving around them. Another important part of this reframing is learning to be okay with not knowing. A wide distribution doesn’t mean your model failed. It means the data is telling you something honest about variability, noise, or lack of information. Once you stop treating uncertainty as a flaw, it becomes a source of insight.
Source: Sahir Maharaj (https://sahirmaharaj.com)
You also become much more aware of assumptions. Bayesian modeling forces you to confront what you believe before seeing the data. That can feel uncomfortable at first, especially if you’re used to models that hide assumptions behind equations. But over time, that transparency becomes a strength. Eventually, uncertainty stops feeling like something you need to “fix” and starts feeling like something you need to understand. That’s usually the point where Bayesian thinking begins to stick. Having gone through this mental shift, the next step is understanding how a tool like PyMC3 actually supports this way of thinking.
import numpy as np
import matplotlib.pyplot as plt
np.random.seed(42)
observations = np.random.normal(loc=50, scale=10, size=40)
prior_mean = 45
prior_variance = 20**2
sample_mean = np.mean(observations)
sample_variance = np.var(observations)
n = len(observations)
posterior_variance = 1 / ((1 / prior_variance) + (n / sample_variance))
posterior_mean = posterior_variance * (
(prior_mean / prior_variance) + (n * sample_mean / sample_variance)
)
posterior_samples = np.random.normal(
loc=posterior_mean,
scale=np.sqrt(posterior_variance),
size=5000
)
plt.figure(figsize=(10, 6))
plt.hist(posterior_samples, bins=40, density=True, alpha=0.7)
plt.axvline(sample_mean, linestyle="--", linewidth=2)
plt.title("Posterior Distribution of the Mean vs Sample Mean")
plt.xlabel("Mean Value")
plt.ylabel("Density")
plt.show()
As a data scientist, the first thing I noticed when working with PyMC3 wasn’t technical complexity, but how intentional it felt. You can’t rush through a model without thinking. The library nudges you to slow down and describe how you believe the data was generated, step by step. Instead of jumping straight to fitting, you start by defining beliefs. What does a reasonable parameter look like before seeing data? How much uncertainty feels realistic? I’ve found that answering these questions upfront often reveals gaps in understanding that would otherwise go unnoticed.
Another thing PyMC3 does well is keep uncertainty visible throughout the process. Results don’t come back as a single value that begs to be over-interpreted. They come back as distributions that invite exploration. You start asking different questions, like how likely one outcome is compared to another, rather than which one is “correct.” From a communication standpoint, this changes everything. When you show distributions instead of points, conversations naturally shift toward risk, trade-offs, and confidence. In my experience, that aligns far better with how decisions are actually made in organizations.
import numpy as np
import matplotlib.pyplot as plt
np.random.seed(123)
x = np.linspace(0, 10, 60)
true_intercept = 5
true_slope = 2.2
noise_std = 3
y = true_intercept + true_slope * x + np.random.normal(0, noise_std, size=len(x))
X = np.column_stack([np.ones(len(x)), x])
prior_mean = np.array([0.0, 0.0])
prior_cov = np.diag([10**2, 5**2])
sigma2 = noise_std ** 2
likelihood_cov = sigma2 * np.eye(len(x))
posterior_cov = np.linalg.inv(
np.linalg.inv(prior_cov) + X.T @ X / sigma2
)
posterior_mean = posterior_cov @ (
np.linalg.inv(prior_cov) @ prior_mean + X.T @ y / sigma2
)
posterior_samples = np.random.multivariate_normal(
posterior_mean,
posterior_cov,
size=5000
)
intercept_samples = posterior_samples[:, 0]
slope_samples = posterior_samples[:, 1]
plt.figure(figsize=(12, 5))
plt.subplot(1, 2, 1)
plt.hist(intercept_samples, bins=40, density=True, alpha=0.7)
plt.axvline(true_intercept, linestyle="--", linewidth=2)
plt.title("Posterior Distribution: Intercept")
plt.xlabel("Intercept Value")
plt.ylabel("Density")
plt.subplot(1, 2, 2)
plt.hist(slope_samples, bins=40, density=True, alpha=0.7)
plt.axvline(true_slope, linestyle="--", linewidth=2)
plt.title("Posterior Distribution: Slope")
plt.xlabel("Slope Value")
plt.ylabel("Density")
plt.tight_layout()
plt.show()
PyMC3 also scales nicely in terms of complexity without becoming opaque. Hierarchies, group-level effects, and varying uncertainty all feel like natural extensions of the same core ideas. You’re not learning a new way to think for each new problem.
Importantly, none of this replaces your existing skills. Your intuition about data, your understanding of the domain, and your experience with traditional models all still matter. Bayesian modeling simply gives those skills a more expressive language.
Once you’re comfortable with how PyMC3 supports probabilistic thinking, the next question becomes how this actually changes real analytical work.
Though, one place I find that Bayesian modeling really shines is when data is limited or messy, which is more often than we like to admit. Early-stage products, small experiments, and niche segments rarely produce clean datasets. In situations like these, traditional models can feel unstable or overly confident. I’ve personally found Bayesian approaches far more forgiving and realistic in these scenarios. Instead of treating limited data as a problem, Bayesian models contextualize it. They blend what you already know with what you’re observing now. This often results in insights that feel more grounded and easier to trust, especially when explaining results to others.
Source: Sahir Maharaj (https://sahirmaharaj.com)
Decision-making also changes in subtle but important ways. Rather than asking whether an effect exists, you start asking how likely it is to matter. That reframing aligns much better with business questions, which are rarely about certainty and more about risk tolerance. Another benefit is how naturally Bayesian outputs lend themselves to storytelling. Probability ranges and likelihoods are easier to explain than abstract statistical thresholds. From what I’ve seen, people engage more when uncertainty is visualized and discussed openly.
Over time, Bayesian reasoning also changes how you evaluate models. Accuracy is still important, but it’s no longer the only thing that matters. Calibration, robustness, and uncertainty all become part of the conversation. Perhaps the biggest shift is cultural. Bayesian modeling encourages humility. Models stop being treated as final answers and start being treated as informed guides. That mindset leads to healthier discussions and better decisions. Having walked through how this thinking applies in practice, it’s worth stepping back and reflecting on what this journey offers you personally.
import numpy as np
import matplotlib.pyplot as plt
np.random.seed(7)
days = np.arange(1, 31)
true_rate = 3.5
events = np.random.poisson(true_rate, size=len(days))
prior_alpha = 2.0
prior_beta = 1.0
posterior_alpha = prior_alpha + np.sum(events)
posterior_beta = prior_beta + len(events)
posterior_rate_samples = np.random.gamma(
shape=posterior_alpha,
scale=1 / posterior_beta,
size=5000
)
posterior_predictive_samples = np.random.poisson(
lam=posterior_rate_samples,
size=5000
)
plt.figure(figsize=(10, 6))
plt.hist(posterior_predictive_samples, bins=30, density=True, alpha=0.7)
plt.axvline(np.mean(events), linestyle="--", linewidth=2)
plt.title("Posterior Predictive Distribution vs Observed Average")
plt.xlabel("Expected Event Count")
plt.ylabel("Density")
plt.show()
And the best part is you don’t need to overhaul everything you do to get started. In fact, you shouldn’t. The best way in is to start small, almost experimentally. Take a problem you already understand well and try reframing it through a Bayesian lens. When the problem feels familiar, the new way of thinking clicks faster. It’s also okay if it feels slow at first. Bayesian modeling naturally asks you to pause, think, and be deliberate about assumptions. That slowdown isn’t friction, it’s clarity forming. Over time, that extra thought upfront often saves you confusion and rework later. One thing I’ve seen again and again is how this approach improves conversations, not just models. When you can talk in terms of likelihoods, ranges, and confidence, discussions become more grounded.
Also, don’t worry about getting everything “right” the first time. Priors won’t be perfect. Assumptions will evolve. That’s part of the process, not a failure. Bayesian thinking is iterative by nature, and your understanding will sharpen as you use it. And if you ever feel stuck, unsure, or just want to sanity-check your thinking, I genuinely mean this: I’m here. Ask questions. Talk through ideas. Sometimes a short conversation is all it takes for things to click. Once that shift happens, Bayesian thinking stops being something you “try” and starts becoming part of how you naturally work with data!
Thanks for taking the time to read my post! I’d love to hear what you think and connect with you
- Kaggle
- Topmate (Free Power BI / Data Science Sessions and Resources)
- Website
- The Tech Journal (Blog)
About the author
Source: Sahir Maharaj
Sahir Maharaj is a Lead Data Scientist who leads the design and deployment of end-to-end AI solutions that drive strategic decisions at scale. As a Microsoft MVP, he has been featured internationally, including in The Indian Express and on New York's Times Square billboards, and is a prolific content creator on LinkedIn.