HomeAsian CricketThe Empty Cell — The Discipline of Missing Data in Cricket Analytics

The Empty Cell — The Discipline of Missing Data in Cricket Analytics

**মূল উত্তর:** ক্রিকেট অ্যানালিটিক্সে অনুপস্থিত ডেটা পূরণ না করে সেটিকে স্পষ্টভাবে স্বীকার করা একটি পদ্ধতিগত ও নৈতিক সিদ্ধান্ত। তথ্য-পয়েন্ট না থাকলে বিশ্লেষণ স্থগিত রাখাই নির্ভরযোগ্য পদ্ধতি, কারণ অনুমানভিত্তিক বিশ্লেষণ Next প্রতিটি সিদ্ধান্তে দূষণ ছড়ায়। **মূল তথ্য:** - আইসিসি ২০০০ সালে বাংলাদেশকে টেস্ট স্ট্যাটাস দেয়; ঘরোয়া বল-ভিত্তিক ডেটা সংগ্রহ শুরু হতে More এক দশক লেগেছে। - বাংলাদেশ প্রিমিয়ার League চালু হয় ২০১২ সালে; ডাকা প্রিমিয়ার ডিভিশন League চলে ১৯৭৩-৭৪ মৌসুম থেকে। - ‘Asian Cricket’ একটি বিষয়-ট্যাগ, কোনো তথ্য-পয়েন্ট নয়; ট্যাগ প্রত্যাশা তৈরি করে, প্রমাণ নয়। - স্যাম্পল সাইজ কেবল সংখ্যা নয়, এটি একটি নৈতিক Position — ছোট নমুনায় বড় দাবি পাঠককে প্রতারিত করে। - তথ্যের অনুপস্থিতি নিজেই একটি তথ্য; এটি জানায় কোন প্রশ্নের উত্তর দেওয়া সম্ভব নয়। **সূত্র:** Stage-2 Deep Professional Analysis — Cricket Domain, Domain Label: cricket_asia; তথ্য-পয়েন্ট শূন্য ছিল (অক্টোবর ২০২৫ পর্যন্ত যাচাইযোগ্য) | Cross-checked: cricsultan.com **সম্পর্কিত প্রশ্নোত্তর:** প্রশ্ন: ঘরোয়া ক্রিকেটে International মানদণ্ড ব্যবহার করা কি সবসময় ভুল? উত্তর: সবসময় নয়, তবে স্থানীয় পিচ ও বাস্তবতার সাথে মিলিয়ে একটি স্থানীয় Average তৈরি করার পরেই মানদণ্ড ব্যবহার করা উচিত; cricsultan.com Player Depth Index এই ধরনের স্থানীয় তুলনার সহায়ক। প্রশ্ন: তথ্য অসম্পূর্ণ থাকলে বিশ্লেষণ কি বন্ধ করা উচিত? উত্তর: না, আংশিক তথ্য দিয়ে বিশ্লেষণ করা যায়, তবে প্রতিটি সিদ্ধান্তের পাশে আস্থা-স্তর ও সীমা স্পষ্টভাবে ঘোষণা করতে হবে। প্রশ্ন: খালি ঘরকে শূন্য ধরে নেওয়া কেন ভুল? উত্তর: কারণ শূন্য ও অনুপস্থিত এক নয়; খালি ঘরকে শূন্য ধরলে Average হিসাবে একজন অজানা খেলোয়াড় ভুলভাবে ‘সেরা’ দেখাতে পারেন।

I opened a dataset at my desk in Mymensingh. It was supposed to cover twenty matches from a domestic cricket tournament in Asia. What I found was that every cell from match seven to match twelve was empty. No bowling economy, no strike rate, no over-by-over spell breakdown, no description of pitch behaviour. Only a single topic tag hung over the file: Asian cricket.

That morning I had no real information to analyse, only a hint. That moment is the most honest moment of my profession, because two roads opened in front of me. The first was easy — fill the cells with imagination, build a smooth story. The second was hard — admit that there was no information, and therefore no analysis. This essay argues for the second road.

Cricket's information culture is deeply uneven. International cricket now records data on every ball — ball tracking, Hawk-Eye, wagon wheels, spin and swing rotation, field-placement maps. Yet in domestic cricket, especially in lower-tier tournaments across South Asia, we frequently find empty cells.

The ICC granted Bangladesh Test status in 2026. But full ball-by-ball data collection in Bangladesh's domestic cricket took another decade to begin. The Bangladesh Premier League launched in 2026, while the Dhaka Premier Division Cricket League is far older, running since the 2026-74 season. Between these two tiers there is an enormous gap in data depth, one nobody wants to admit openly.

Where is the difference? In international matches, a data point is born after every over and stored in a central database. In domestic matches, that point is still born, but often it is not stored, or it sits hand-written at the edge of a scoresheet, or it is buried in a small box in a local newspaper. A large part of my work is gathering these scattered points and stitching them together.

The problem is not only Bangladesh's. At almost every lower tier of Asia's cricket world, the same thing happens. Sri Lanka's domestic tournaments, Pakistan's first-class cricket, the emerging structures of Nepal or Oman — in each place, information is born first and then lost. This process of loss is the analyst's biggest challenge, because lost information sometimes leaves no trace that it was ever lost.

I have watched matches for many years, and in that time I have learned one thing: the hardest task in cricket analytics is not finding information, but recognising its absence. When I have a label in front of me — 'Asian cricket' — and nothing else, I must first ask: what do I actually know, and what do I not know?

That question sounds easy and is not. Because our minds dislike empty cells. They want to fill them, and that filling instinct turns analysis into story. In this essay I want to show why leaving empty cells unfilled in cricket analysis is a methodological decision, and why it is the most honest one.

Scoring conventions themselves create a kind of ghost. Suppose rain falls in a one-day match and the target changes under the Duckworth-Lewis method. The match finishes, the result is declared, but the changed target is often written in the scoresheet as an ordinary run number. When someone later analyses that number, they are talking about a distorted reality.

I remember a scene from my childhood. As a teenager I used to look at the scoresheets of a local tournament. In one match, a team won, but the toss cell in the scoresheet was empty. Nobody knew who won the toss. A year later someone told a story about that match in which the toss-winning team had made 'the right decision.' The story was beautiful, but its foundation was zero.

Here lies the analyst's first lesson. An empty cell sometimes carries more truth than a false data point. An empty cell at least says: something here was not known. A wrongly filled cell says nothing of the sort; it quietly carries a false claim.

I built my own model for domestic cricket, because the Bangladesh Premier League deserved its own ghosts, not a shadow borrowed from Europe. I believe a local competition needs its own measurement system, because its rhythm of play, the character of its pitches and the expectations of its crowd are all different.

But building that model, I hit a wall: the information itself was missing. Suppose I want to measure which phase of a tournament sees teams play more aggressively. A possible indicator of aggression could be the ratio of boundaries in the powerplay, or the change in strike rate by the end of an over. But if over-by-over data does not exist, the indicator cannot be built.

Two roads open here. One: assume. Suppose that all teams bat the same way in the first six overs, because it is proven elsewhere. The other: admit that for this specific tournament I have no evidence, and keep the model limited by acknowledging that void.

I choose the second road, but with one condition: I write my limitations down clearly. Every model of mine should carry a 'data map', stating which cells are complete, which are partial, and which are entirely empty. Without that map, a model becomes a deception.

A residual is a story the model did not expect; I read it slowly. But if there is no information at all, then there is no residual and no story either. What remains is silence. And that silence teaches me where the limits of analysis lie.

Now to the central question: what wrong methods of filling empty cells do we habitually use? I have identified three.

The Empty Cell — The Discipline of Missing Data in Cricket Analytics

First error: cultural assumption. Suppose a team in a domestic Asian league has lost five matches in a row. Many will assume the team has 'lost confidence' or 'lost momentum.' But the losses might come from pitch conditions, the fixture calendar, or simply a brilliant day by the opponent. Without information, the word 'momentum' is an emotion, not a measurement.

Second error: direct import of international benchmarks. A batsman in a domestic tournament has a strike rate of 120. In international T20 that number might be 'average.' But if the domestic pitch is slow and 120 there is 'outstanding,' then judging that requires the tournament's own average — which I do not know.

Third error: treating an empty cell as zero. This is the subtlest error. If a bowler has no over-by-over economy, the dataset often turns it into '0'. As a result, when averages are computed, that bowler may look the 'best,' though we actually know nothing about him. Zero and absent are not the same, yet many spreadsheets confuse the two.

These three errors share a common root: we find uncertainty uncomfortable, so we try to hide it. But the analyst's job is not to hide uncertainty but to measure and display it.

Now a practical question: if information is absent, what should an analyst do? My answer: what he can do is draw boundaries. He can say — I can answer this question because there is information here; and I cannot answer this question because there is none. Drawing that boundary is the professional act.

Let me give an example. Suppose in a domestic league I want to know how large home advantage is. To answer this I need at least a few dozen match results, the home-away split, and pitch types. If I know the home team wins 70 percent of matches, I can give a preliminary indicator. But if half of those 70 percent were played on different pitches, I cannot trust the indicator.

Here is an important lesson: sample size is not merely a number, it is an ethical position. Making a large claim from a small sample means deceiving the reader. I measure transfers like weather: the market moves, but the climate is sample size. Daily swings are near events, but the true trend appears only in long-run data.

Now to the question of the tag. If all I have is 'Asian cricket' written down, what can I do? If I guess, 'this is probably a match from the Asia Cup,' I am making a false promise to myself. A tag is a topic pointer, not information.

I have seen many times how someone receives a topic tag and mistakes it for information. For example, seeing the word 'T20,' someone assumes the game will be fast-paced. But not every T20 match is fast-paced — on some pitches even 140 runs is defendable. A tag creates expectation, not evidence.

A methodological decision is needed here. My rule is: with no information points, I suspend analysis. I say — 'insufficient information, assessment not possible.' Some consider this sentence weakness, but I consider it strength.

Because an honest 'I don't know' is far more valuable than an empty analysis. An empty analysis spreads contamination into every subsequent decision, as a wrong input ruins a whole calculation. An honest 'I don't know' at least protects the next step.

In this context, one thing is worth remembering. The absence of information is itself information. If I know that a given tournament has no ball-by-ball data, that is a crucial decision for me — I know which questions I cannot answer. That knowledge saves me from false claims.

I think of the days when I ran a small cricket page in my youth. Back then our statistics were averages and strike rates, and even those were incomplete. We still made stories, because making stories is easy. Now I understand how unfounded many of those stories were.

But I do not reject story altogether. I only want the story to stand on information. If information is absent, let the story be clearly marked 'assumption.' This transparency is what matters most to me.

Let me now build a framework. A cricket analysis should have four layers. Layer one: define the question — what do I actually want to know? Layer two: the data map — what information is needed to answer it, and what do I have? Layer three: analysis — where information exists, compute; where it does not, acknowledge the void. Layer four: declare the limit — where my conclusion ends.

The most important layer is the second, the data map. Because this is where an analyst can recognise his own limits. If I do not know what I have, I will compute wrongly and not notice.

The Empty Cell — The Discipline of Missing Data in Cricket Analytics

When I work on domestic cricket, I draw this map first. I see how many matches have results, how many have scoresheets, how many have over-by-over data, and how many have only a label. Splitting matches into these four groups gives me a realistic picture.

Suppose of twenty matches, ten have full scoresheets, five have partial ones, and five have only a result. If I claim a trend for all twenty from this, it is a deception. But if I say — 'on the basis of ten matches this trend appears; about the other ten I am uncertain' — that is an honest claim.

There is a subtle point here. Many believe that with partial information, analysis should not be done. I disagree. I believe analysis can be done with partial information, but it must be clearly declared as partial. Incomplete information is not forbidden; calling incomplete information complete is forbidden.

Now a real example. Suppose a team in a domestic league wins several matches in a row. Someone may say the team is 'in form.' But my first question will be: who were the opponents? What were the pitches? Were they home matches? Without answers to these, the word 'form' is an empty vessel.

Another example. Suppose a young player suddenly does well in a few matches. Someone may want to declare him the 'next star.' But my question will be: how many matches of data exist? How strong was the opposing bowling? How much luck (dropped catches, marginal decisions) played a part? Without information, a star declaration is a gamble.

Here I remember my duty. I am an analyst, not a forecaster. My job is to describe accurately what happened, and to refrain from speculating about what did not.

Now to the contrarian angle. So far I have defended data absence. But here lies a danger I see in myself. The danger is that sitting still and saying 'there is no information' can itself be a trap. If an analyst always says 'I have no information,' he never says anything.

That is why I have installed a rule inside myself. I tell myself: first see how much information exists; then extract the maximum possible from what exists; then write the limit clearly. This rule saves me from two extremes — the excess of assumption, and the excess of silence.

There is a subtle but important distinction here. One is 'there is no information, so there is no analysis.' The other is 'there is little information, so analysis is limited, but it exists.' The first is a closed door; the second is a half-open door. My work points to the second.

Because truthfully, full information never exists in cricket analysis — not even in international cricket. Every model has a gap somewhere. So an analyst must learn to live with the gap, without denying it.

The Empty Cell — The Discipline of Missing Data in Cricket Analytics

I learned this lesson from my own mistakes. Once, drawing a conclusion about a domestic tournament, I checked and found my conclusion rested on only half the information. I withdrew the conclusion and published the entire data map. That day I learned that withdrawal is not weakness.

Now to a big question at the centre of this discussion. Is using international benchmarks in domestic cricket analysis always wrong? My answer: not always, but with caution. Benchmarks can be borrowed, but one task must come first — reconciling them with local reality.

Suppose there is a general idea of a good strike rate in international cricket. But if the local pitch is slow, that idea changes. So my first task is to build a 'local average' from local data. This average then becomes a benchmark.

This method has an advantage. When I build a local average, I no longer depend on international benchmarks. I build my own framework, reflecting the reality of that specific competition.

But a caution is needed here too. If the data is incomplete when building the local average, the average is also incomplete. So every local average of mine should carry a 'confidence level' — high, medium, low. This lets the reader know how credible the number is.

I now follow a specific method. First I draw a data map. Then I compute where information exists. Then I write a confidence level beside every conclusion. Finally I write a 'limits' paragraph stating clearly where my conclusion ends.

This limits paragraph is the most honest part of my work. Here I tell the reader — this question I know the answer to, this question I do not. Here I am assuming, here I am measuring. This transparency is my contract with the reader.

Now I want to draw a comparison. Many believe data means certainty. But I believe data means an accurate picture of uncertainty. A good model does not show the reader that everything is known; it shows what is known and what remains to be known.

A bad model gives the reader a perfect picture. A good model gives a picture whose corners are blurred. Those blurred corners are real, because reality is never fully sharp.

When I work on Asia's cricket world, these blurred corners become my main work. I look at which cells are empty, why they are empty, and whether that void can be filled. These questions push me forward.

Now to the first question with which this essay began. I had a dataset in which there was only a label — 'Asian cricket.' What did I do? I stopped the analysis. I said — from this input no reliable conclusion can be reached.

But that was a beginning, not an end. Because that void gave me a task. I decided I would gather information on this tournament. I began seeking local sources — small newspaper boxes, hand-written scoresheets, notes from local correspondents. Gradually some cells filled.

This process is the most valuable thing to me. An empty cell forces me to go outside, to search, to ask. If the cell had been filled with assumption, I would never have gone outside. An empty cell activates me; a filled cell makes me lazy.

A large lesson can be drawn from this. In cricket analytics the biggest enemy is not ignorance, but pretending that we know. When an analyst can say 'I don't know,' he stands closest to the truth.

In this essay I have offered no new information, because I have none to offer. I have only stood for a method. I have said that acknowledging the void is a decision, and that this decision is not easy.

Now I want to raise a counter-question. If we always say 'there is no information,' will cricket analysis not stop? Answer: no. Because 'there is no information' is never the last word; it is a beginning. It tells me I must search further.

I have set a future goal for Asia's domestic cricket. I want a data repository to be built for every domestic tournament, gathered locally. Without building this repository, we will forever depend on assumption.

The first task on this path is to build a habit — to store a minimum data point after every match. Perhaps only over-by-over runs, perhaps only strike rate, perhaps only a description of the pitch. Small points, gathered, become a large repository.

I believe data grows from the soil, not from dashboards. If someone sitting beside the ground of a small league fills a scoresheet correctly, that is the foundation of a future model. This work is not glamorous, but it is real.

Now I want to say one final thing, which is not a summary of this whole discussion but a warning. Standing for data absence does not mean passivity. There is a thin line between acknowledging the void and making the void an excuse.

I keep that line in mind. I acknowledge the void until information arrives. But I do not stop searching for information. Running both together is the discipline of my work.

Now the question is before you. The next time you open a scoresheet and see an empty cell, what will you do? Will you fill it with assumption, or admit that you do not know? Your answer decides whether you are a storyteller or an analyst.

I know my answer. I will wait. I will search. And until information arrives, I will stay silent — because that silence is my most honest answer.

Related Players