Get to know our data detectives – Ivo

data detective Ivo

Come and meet Ivo, one of our data detectives and our data scientist. Venture into his curious mind and discover what makes him love his job, what pushes him forward, and what is his outlook on the world. As one of our talented blog post writers, this interview came easy to him, if the inspirational answers are anything to go by. Tell us who you are and what you do in Digital Poirots. I’ve been asking questions about nearly everything since I was a child, and that is something that remains an integral part of my personality to this day. I’m very curious about the world that surrounds me and I find great joy in learning new things. I can never get tired of asking questions and exploring various things from a number of different areas. As for what I do, when I’m not jogging, listening to an interesting podcast, enjoying my favorite music, watching a classic film, reading a book, going to the theater, answering questions on a quiz, or traveling, I’m working as a Data Scientist with Digital Poirots. What does your typical day look like? In order to get into the character of a data detective, I like to start off my day with a nice cup of tea, just like a lot of famous fictional detectives from Sherlock Holmes to Patrick Jane do. Afterward, I try to use my skills, mostly in Python and SQL, to get to the bottom of data-related challenges and problems. To get myself focused and relaxed, I like to listen to music while thinking about and solving the cases at hand. When I’m not using my skills, I’m doing my best at improving them because I’m aware that each day brings new opportunities to learn something new and to grow, both personally and professionally. What made you decide to develop your career in this field? Given that I believe knowledge is the power that should run and shape our society, I feel being a Data Scientist is a perfect fit for me. This is a job that lets me use my creativity and intuition to try to extract knowledge from an endless wasteland that data can sometimes be. What drives you in your work? At times I like to think of my job as an everyday opportunity to reach my own eureka moments that make me feel fulfilled. It really gets your blood flowing when you experience one of those famous aha moments where the light bulb shines above your head. What’s the best part about your position? I look at every new project as a new adventure I’m departing on. That’s because I never know where the data will take me and what conclusion I will come to which is pretty exciting from my point of view. Furthermore, I get to work on projects from various domains which keeps my job from getting monotonous and repetitive. All this makes me feel free and we all know how nice and sweet freedom can be. What don’t you like about your job? I often say that technology is the best thing that has ever happened to society while at the same time being the worst thing that has ever happened to society. It completely changed the way we live and work and, frankly, I’m not happy we spend so much time using technological devices, both at work and in our spare time. I believe we all need more interaction with each other and that technology kind of alienates us. I feel it would be better if jobs in the IT sector involved more one-on-one interaction with other people. What’s the biggest mistake you’ve made in your career or what ups moment you had? I can’t remember a specific big mistake I’ve made during my career, but there are always minor mistakes each of us makes on a daily basis. For example, knowing many programming languages can cause confusion between keywords from one language with keywords from another language. It has happened to me a couple of times and, while it was sometimes really frustrating, it can get amusing, as well. The most important thing is to be self-critical when necessary and to be ready to learn from mistakes. What drives you crazy about your job or your daily activities? There are a few things I’m really annoyed by. For starters, traffic jams are daily situations where I feel I’m wasting my life just sitting in a car, a streetcar, or a bus and that’s when I get very frustrated. Furthermore, regarding my daily routine, I don’t like doing anything repetitive. I believe the main purpose of computers is to relieve us of doing repetitive tasks over and over again and increase the quality of our lives by doing that. Actually, that’s one of the main reasons why I appreciate their existence. Which technology or tech stack do you like the most? I don’t have a problem working with any technology necessary for the given project. Even in the ones I’m not yet familiar with, I’m eager to learn whenever and whatever is required. However, I feel the most confident when working in Python, R, Java, JavaScript, and SQL because these are the technologies I used the most to this point. If I had to pick one among these mentioned technologies, it would definitely be Python with its data science stack that offers a wide range of possibilities and comes with a very active community. What advice would you give to someone entering this field? You don’t need to have a lot of experience or knowledge to assume that the most important thing in this field is to understand the basics very well. When you know and truly understand the basics, it’s much easier to improve yourself afterward. Also, be proactive – ask the questions and seek the answers on your own, that’s the best and the most natural way to grow. Do you have any funny or interesting stories that happened

All you need to know about data misinterpretation and misuse

data misinterpretation

Data misinterpretation comes from deliberate or unintentional false data interpretation. Often different data users will interpret data differently. It depends on the level of data literacy and what exactly users need for the data to show. Various influences determine how we see and use data, and wrong representation can lead to incorrect insights and metrics. Wrong or incorrect data interpretation can occur when integrating data democratization into business operations. This can lead to making wrong data-based decisions and must be avoided at all costs. But, it can prove to be difficult to minimize the influence of bad data interpretation. It’s hard to avoid it on the individual level. It has to be rooted deep down in individual users in how data should be approached and handled. If a user, an employee in most cases, isn’t educated in data management and analysis, then it’s more likely for them to make mistakes when dealing with data. What‘s the cause of data misinterpretation? When looking at what causes users to misinterpret or misrepresent data, there are multiple reasons. Some can be credited to the lack of data literacy and some to psychological aspects. All of these factors influence how data science and data engineering results are perceived, and it’s something that has to be taken into account when jumping into data-based projects. Inadequate data Data comes in all forms and not all of it can be analyzed the same way. There are processes needed to be performed before any data is ready for any sort of analysis. Actions like data cleansing, deduplication, identification of bad data, and such precede future methods for data analysis. This is where proper data management systems matter so you can avoid working with data that isn’t ready for data science processes or any other form of analysis, for that matter. Insufficient or unrepresentative data Anyone who had any interactions with statistics knows that small samples or insufficient data lead to statistical insignificance and wrong conclusions. Sample size or dataset size determines if the results are relatable to what we observe. Looking at too small sets of data can often skew our findings and present results that are not representative in any way. Not understanding data Data literacy probably plays the biggest role in data interpretation. Those that do not understand data, how it’s collected, what it presents, or how to analyze it, will draw wrong conclusions. Before any data analysis actions, users must understand what the data presents. They need to know why something was collected or generated, what are the data sources, what’s the format and what each piece of information presents. If users don’t understand the domain, they have difficulties understanding metrics and insights. Lack of context If users aren’t familiar with the reasoning behind data collection and analysis and are presented with just the final results, they’ll have issues with fully understanding what’s going on. How people interpret results depends on how they analyze and contextualize data. Having access to all data doesn’t mean that users should always interpret it without knowing the source, meaning, and intent. Attribution bias and lack of information Often, users make snap judgments with limited information. Without deep understanding and knowledge, users make decisions that haven’t taken all the data into account. They make presumptions without the full context. If certain data or variables are omitted, one cannot draw the correct conclusions and it will ultimately lead to faulty decisions. Faults in aggregation Data is usually aggregated and observed as a whole set when in most cases it should be observed on more detailed levels. Huge samples of data can lead to the observance of information that can lead to wrong conclusions. Sometimes it is necessary to fragment data to see conclusive patterns or insights. For example, if we view customers or clients as a whole, then we will miss the distinction between segments and different profiles. And each segment will provide different insights and characteristics that can lead to more specific campaigns and actions. Oversimplifications of findings If users get some metric or insight based on data, often they see it as it is. They interpret it as a simple result, rather than including all the reasons and influences behind that result. Imagine using multiple strategies for advertising, and you explain the increase in sales based on just one strategy or advertising channel. Omitting other influences and doings and discarding detailed analysis can cause wrong data interpretation. Correlation doesn’t equal causation Just because some events are correlated doesn’t mean that one causes the other. Some data can seem to correlate, meaning that one thing stimulated the other, but that’s not always the case. A change in one variable does not automatically induce a change in the other variable. Misleading data As often some might misinterpret data, using certain visualizations can also mislead how users perceive information. Choosing one way of visualizing data can distort metrics and change how the viewer understands those figures. Sometimes, misleading visualizations are created intentionally and sometimes they are not if users don’t think it through when choosing visualizations.The causes of misleading data visualizations range from wrong visuals and information manipulation to unclear data sources. Manipulation of axis and scale One of the most used methods of misleading visualizations is through axis and scale manipulation. Usually, the axis or scale begins with a zero, but if the intention is to overdramatize results, the scale can start from a greater number and the range between numbers on the scale can be larger or smaller depending on how someone might want to show their data. This is mostly used in bar charts, where depending on the scale, bars can be made to appear taller or smaller. Data obscuring Data presenters often obscure results by omitting information that doesn’t serve them. Some might rather point out the positive or negative, depending on their agenda and goal. To achieve that, the data that doesn’t support the objective will be removed from visualizations. This way the focus is on

Deep dive into data democratization – the good and the bad

Data democratization

Data democratization is being labeled as one of the most attractive trends in data science and data management for the upcoming years. It’s been raising hype for quite some time and it’s not slowing down. As companies face more and more data, they need effective ways of handling it, but at the same time, they need to give their employees access to it. Long gone are the days when the IT department was the only one with access to important data. With how fast the market and business environments change, on-time responses and actions are vital to staying on top of things. So what is data democratization? Simply put, it means that everybody has the access to data they need, relevant to their role in the company. Its main purpose is to allow users, or employees, in this case, to use data for drawing insights and metrics much needed to form decisions. Data democratization implies that users need to be educated on how to work with data and that data literacy is a top priority. What’s important is to cut out the middleman, the IT, so information can move faster and new opportunities can be spotted more easily. This is where self-service analytics come into play. Why is it a good thing? There are vast benefits when it comes to data democratization and each one meets a rise in importance. Employees empowerment & better collaboration By giving employees, on all levels, access to data, companies empower them to make decisions that matter. They get a sense of a collaborative environment that focuses on the team. If employees can analyze and use data on their own, they can upgrade their skills and communicate better with others across different departments and teams. It enables knowledge sharing among employees and drives innovation. With more freedom in data handling, employees get a greater sense of purpose. Faster decision-making If employees don’t have to go through the IT department each time they need access to some data, they can save time immensely. If users get instant access to insights and metrics, they can make decisions faster and more accurately. Information flow is more efficient and reaction time is reduced. Lower costs Faster decision-making and cutting out the hierarchy in data accessibility, lower cost in the long term. Business operations and people are more effective, and with data in their hands, they more easily spot areas where costs can be reduced and revenue increased. Flexibility Data democratization provides more agility to business operations. By spreading data through all hierarchical levels, cooperation is more fluid. Reactions to changes are quicker and users are more adaptable to certain micro and macro influences since they don’t have to wait to analyze much-needed data or information. There is no rigorous structure regarding data access and teams can be more agile in their work. Increased effectiveness If more people have the skills and knowledge to analyze data, there are more eyes that can spot opportunities or threats, for that matter. Faster data accessibility means that certain tasks are performed more quickly. Employees perform better since they have information at their fingertips. Certain business operations can be sped up and done better. Amplified data-driven culture Data literacy is an important aspect of data democratization. If all employees get skilled in data management and analysis, it can promote data-driven culture and performance based on data. If all employees understand data and how to utilize it, then they’ll more confidently explore data and convey findings between departments or across teams. When data democratization goes wrong Even though data democratization has a lot of benefits, there are always some downfalls. Especially, if the company hasn’t prepared itself correctly for data democratization implementation. There will always be a level of mistrust in giving everyone access to data. And if not implemented correctly, it can cause quite a stir in the company. Data privacy If everyone gets access to data, then how can data privacy be enforced? What if someone unintentionally misuses data and breaks privacy regulations and laws? This is a valid issue because we all know mistakes happen. Data leaks can slow down or even endanger business operations. Data silos Data silos are still an issue when data is not centralized or organized in a way that it can’t be duplicated or found across various databases and in different formats. If the initial setup doesn’t support a single source of data, then this can prove to be complicated and difficult in the long run. If each department has or creates its own source of data and isolates it from the rest of the organization, then data democratization can’t be integrated fully. Employee trustworthiness There has to be a certain level of trust that employees or users won’t misuse data to their or someone else’s advantage. Companies generate or collect sensitive data that has to be handled carefully and in line with regulations. Fraud and risky activities can easily damage not only private data and business, but their reputation as well. Misrepresentation or misinterpretation of data One of the biggest challenges in data democratization is that different users will differently interpret data and the information it provides. Often the same piece of information can be shown in different graphs like in the example below. It’s the same data, the only difference is in how an axis is numbered or labeled. A different interpretation caused by different representations can lead to wrong conclusions and insights. Misinterpretations can lead to instability, costs, and loss if the decision, based on inaccurately presented data, was the wrong one. Image 1: Example of the same data visualized in different graphs Cost of democratizing data Some might get discouraged in implementing data democratization because of the initial cost. You have to educate your employees about data, how to utilize it, and how to work with analytics tools. Amplifying data literacy comes at a cost and initial technological infrastructure comes at a price as well. Conclusion Data democratization is a continuous process

How to determine KPIs for a Retail BI Platform

How we determined KPIs for our Retail BI Platform

When talking about KPIs we must understand their applicability and what they actually mean. In any data science and data engineering project, the determination of KPIs, or rather metrics and insights, is an important step. How would you know what to measure and what can you actually do with all that data if you don’t define which values you need and want to show? When working on our retail business intelligence platform, we first needed to understand the industry and data itself. Which results are important for retail? How do they measure success? Which data is actually available to us? In the end, each value represented something, and we needed to analyze and understand what that something is. How to approach data It was easy seeing data for what it is. A series of values over time. And when you look at certain values, they represent an action. In our case, we had sales quantity, price per SKU, location, date, product ID, department, and product category. The data set told a story of retail traffic and consumer spending. If observed as individual values, they can be easily interpreted. But, we wanted to create a solution that will draw extra benefits and explanations derived from the original data. Our aim was to use this data set to bring value to more than one department or cost and profit center. We didn’t want to see data through a singular explanation. Certain numbers can speak to a larger variety of people and departments. One metric can influence decisions and actions in more business operations and procedures. Since the dataset came originally from a data challenge, our task was to estimate unit sales. So, if we have retail data and SKU analysis, we needed to start from that. The question we asked ourselves was what is important in a simple SKU analysis? What does our data tell about SKU performance? And from our data set, the obvious answer was sales. What is the level of revenue per SKU is a pretty straightforward metric. We can now observe which SKUs are better performing and what level of revenue they can bring to a retail store or location. KPIs through multiple variables and dimensions But, this one metric won’t suffice. We needed to observe values across more than one attribute or variable. And we didn’t want to focus just on static data. Our goal was to be able to predict future SKU movements. So, it was time to develop a matrix where we could distribute our data to outline basic metrics our solutions should show and continuously calculate. Table 1: KPIs through multiple variables and dimensions Individual SKUs, if observed across time and locations, tell a variety of stories. For example, one product can have high sales in one location and low in another. This tells us that there are possibly different segments of consumers. They can differentiate based on income, culture, age, consumer preferences, upbringing, etc. But it can also be an indication of different marketing and sales effort in this certain store or area. Or even perhaps, this SKU has varied in stock in different stores. There could’ve been backlog or overstocking. Maybe there was a new trend in that location that stimulated the attractivity of that particular product. So, as we can see, one metric serves multiple explanations and applications. And that’s why this matrix was developed. These SKU KPIs can be compared across multiple variables or dimensions. This is also a basis for the streaming component of this software solution. These metrics needed to be available to users in real-time, so they can access them anytime for their reporting and instant reactions to unexpected changes. Our objective was to allow users to use data in accordance with their needs, so they wouldn’t need to wait until month-end reporting to see their retail performance. Also, based on those simple metrics and SKU analysis the solution can send real-time notifications and present them to users. Of course, it’s not only values but through visualizations, users can more effectively interpret results and share them with other stakeholders.  What we can take from this approach is that even the simplest metric analyzed through more than one dimension can provide insights and explanations for numerous events. What we must pay attention to are restrictions derived from the chosen technology. We can imagine and define KPIs and metrics, but sometimes data doesn’t support them. And the chosen technology can offer limitations in executing calculations. Not every value can be turned into an impactful KPI. And not every value should be measured. The business requirements have to be aligned with the data science and data engineering side of the process. KPIs that tell a story After the matrix and simple KPIs validation, we turned our focus to more complex KPIs. We wanted to see which KPIs are possible to calculate if we have only certain values. For example, if we have units sold and a price of SKU, what can be calculated from that? And of course, the KPI had to make sense to the end user.  This is why we decided to create another table that outlined our major KPIs across periods of time and location since that was our main filtering determinant. Table 2: Major KPIs groups across different locations and time Revenue indicators Our first KPI is sales or revenue a certain SKU can bring. We observed it across 5 different levels (from product level to cumulative stores level) and 6 time dimensions (from one day to all time or the whole period for which we have data). Sales or revenue KPI is calculated as sold quantity x price per SKU. This KPI answers the question of which SKUs are outperforming or underperforming others. It’s an indication of which ones play a bigger role in sales performance.  When also looking at revenue per unit, we can get so many insights about products sold in the store. We can answer questions like: Market share The next one

Data streaming and how to influence business decisions

data streaming

In data streaming, data flows in continuously. That means that data is processed instantly and in real or near real-time. This is an advantage over batch processing where it’s required for data to be downloaded in batches first. It’s an obvious benefit for decision-making. Here, decision-makers get data at the time of the event and not after the facts. So imagine having multiple data sources, from electronic devices, IoT devices, mobile phones, cloud services, databases, sensor devices, and so on. That’s a lot of data coming in from everywhere. Now imagine if important data from these sources came too late. Your reactions to changes and disruptions are also too late. You have missed important facts that could’ve helped you stay on top of things. What does data streaming do for you? Data streaming is based on dynamic data intake, meaning its processing is done on instantaneously consumed data. While it ingests and collects data streams from multiple sources, it simultaneously stores and aggregates the data for processing. That’s the basic concept of Kafka, a streaming technology that has risen in popularity. We have covered Kafka in one of our blog posts since it’s the main part of our retail business intelligence platform. Companies whose products or services depend on quick responses and correct information have implemented streaming technologies as their core business value. They have built streaming architecture to support not only their daily operations but customer engagement and experience as well. For data consumption in streaming, there are a couple of stages. Stream ingestion, stream processing, stream analytics, and data presentation in selected data topics. Basically, streaming technologies ingest data from various streams or sources, process it through aggregation and transformation of data, turn it into actionable values or insights, and present data topics to users. Why is this valuable? Well, real or near real-time information and data that is processed at such a speed mean less time spent analyzing static data which in some cases loses its value. For example, if you look at data regarding financial markets, you need updated information since it can change instantly. What happened yesterday doesn’t have to be the same for today, and high-stakes decisions rely on real-time data. Each event-streaming application has these basic functions or purposes: stream processing and data integration. In stream processing, it has to have the ability to process or transform consumed data. For data integration, it has to feed these events to other data systems like data warehouses, lakes, etc. Why Kafka and how does it work? Apache Kafka is a widely used technology and it’s extremely popular for event storing and streaming. What Kafka does is it receives and stores messages from producers to a server that’s called a broker. Those records are arranged into topics. Multiple brokers compromise a cluster. The aforementioned topics are served to consumers who subscribe to the ones they want.  So, in short, we have event producers, topics where these events are organized, and consumers.  Kafka is a fast and flexible tool, and it decouples data producers from processors with better latency and scalability. That’s why it’s the first choice for data streaming in most cases. The one very cool thing Kafka offers us is scalability. Furthermore, it does it smoothly because we don’t have to worry about rebalancing after adding a new consumer – Kafka automatically takes care of it. Kafka’s main purpose is to produce and consume messages, without putting emphasis on data processing. On the other hand, Spark is a well-known, easy-to-use, distributed processing engine used mainly with big data. The core idea of Spark is to allow us to efficiently transform and process large amounts of data in real time. Product or service vs business operations Data streaming can be used for business process optimizations, decision-making, and for creating products or services. When we talk about decision-making it is obvious where data plays its role. Users can make decisions and perform actions instantly if they get instant access to real-time data. If data streaming is based on data from business operations, users can spot errors, discrepancies, or opportunities in the processes themselves. But when we talk about products or services, this is where data streaming really comes into play. What matters in business is customer satisfaction and experience. They expect a great service or product every time, and data-based products and services have to be executed perfectly. For example, think about e-commerce and online shopping. One area where data streaming comes into play is stock information. Buyers need information on whether an item is available. Also, data streaming focuses on bringing recommendations to customers based on their searches and purchases so they get offered products up to their tastes. This stimulates consumers to fill their cart and ultimately finish the purchase. Or let’s look at taxi apps. When a user gets on the app, they need information on when the pick-up will be, what’s the price and what’s the estimated time of arrival based on traffic information. Data streaming is what streams that information to form a service. And users or customers depend highly on the accuracy of those pieces of information. Even content streaming services use data streaming technologies to track user activity and present the best offers and recommendations to the end users. The benefit of data streaming as a base for a product or service is the key to increasing customer value and to creating impeccable user experiences. What is the value of streaming technologies? You get multiple benefits by integrating streaming technologies into your data processing operations. They can strengthen your business, accelerate decision-making and anticipate upcoming disruptions or issues. Respond in real-time Streaming data directly at the moment of generation is key. Reducing the time between when an event is recorded and processed presents an opportunity to minimize response lag time. Users get more confident in making decisions, especially if they operate in fast-paced environments and such industries. Time is of the essence and staying behind is not an option.

Get to know our data detectives – Renato

data detective

It’s time for another one of our interviews where we put our data detectives in the spotlight. So, we caught up with Renato, our data team lead. He allowed us to venture into his data science and engineering world. Not only is he the data team lead, but he also plays an important role in business development processes. From his daily activities to work challenges, he covered all the topics for us to get to know him better. Tell us who you are and what you do in Digital Poirots and Deegloo. I perform a couple of roles here in Digital Poirots and Deegloo. My engineering half is focused on data science and engineering, where I perform the role of a team lead. As a co-founder, I have an entirely different role which is related to business development processes. What does your typical day look like? It’s a mix of business and engineering tasks. 🙂 Repetitive tasks are rare as each day brings some new challenges to overcome. Some typical data-related tasks include data collection and processing, development of data pipelines, analysis and visualizations, and development of prediction models. What made you decide to develop your career in this field? My interest in math started during the second half of elementary school and increased in high school where I had a great teacher.That pushed me toward the Faculty of electrical engineering and computing where I found myself in the field of computer science. What drives you in your work? The sense of purpose. It’s a great feeling when you know that the work you did made an impact. What’s the best part about your position? The fact that every day brings some new challenges. What don’t you like about your job? The fact that every day brings some new challenges. 😅 What’s the biggest mistake you’ve made in your career or what ups moment you had? This happened quite a while ago, at the beginning of my career. I was doing maintenance of the ETL process that was implemented with the Pentaho Data Integration tool. There was a task to update the list of emails that receive certain notifications from the ETL pipeline. Pentaho defines the pipeline as a sequence of transformations and jobs, where each one of them is defined as a separate XML file. I checked the files and found out that emails were hardcoded. This meant I could not solve the task just by changing environment variables. As we had a lot of files that required the change, I decided to write a Python script that parses XML and replaces hardcoded emails with a new parameter. Because of a bug in the script, I ended up replacing not only the emails I should have but others as well. Hopefully, we had a backup that I created before script execution, which enabled fast rollback.  In the end, it took me more time to write the script, find the issue, and fix it than it would take me to do all of the changes manually in the first place. Since that day I am way more careful and reserved about applying automated processes to such sensitive tasks. Automation is a good thing, but there are some scenarios where manual intervention is a better option. What drives you crazy about your job or in your daily activities? Too much multitasking. In these situations, it’s best to prioritize tasks and reorganize the work. Otherwise, you usually end up with stress and unfinished business. Which technology or tech stack do you like the most? For data engineering, I mostly combine SQL with ETL development tools. Depending on the project and requirements, ETL pipelines are done through no-code tools like Pentaho Data Integration or code-based technologies like Airflow.  When talking about data science and machine learning, it’s mostly related to Python-based stack and tools like Pandas and scikit-learn. For the visualization part, I mostly use Tableau which enables me to create good-looking, interactive, and dynamic visualizations in no time. All solutions we create are hosted on the AWS cloud, where services like RDS, EC2, S3, and Sagemaker come into play. What advice would you give to someone entering this field? I would recommend learning in a structured way. This is the advice I got from one of my favorite university professors Jan Šnajder, who lectured to me on Machine learning during my studies. He compared the learning process with books. Each day we learn, we read a new book. Once we are done with the lecture, we put the book on the shelf. As time passes, more and more books will be on the shelves, and we would forget the details that are written in them. But when the time comes, and we need some information, we know which book it belongs to. Additionally, we know the relationships between the books, which allows us to group concepts and establish a hierarchy between them. This probably works for other fields as well, but I found it especially important in the fields of data science and machine learning. Do you have any funny or interesting stories that happened here in the company? The funniest thing that happened to me recently is that I almost tipped over from pedal Go Kart during the race we had on our team building. After a few meters of 2-wheel driving, I ran out of the track and crashed into the barriers. I guess it was funnier for the audience than for me. Your favorite person to work with? This is the hardest question here. I can’t decide on one person, because everybody I work with at the company is professional and pleasant to work with. If you can compare your job to one movie or show, what would it be? I like Interstellar very much. Doing our data detective work sometimes reminds me of Murph. If you can choose one song to go along with your job or which would make you be really in the zone while

Get to know our data detectives – Dario

Het to know our data detectives

If you all wondered who our data detectives are and what they are like, well, dive into this short interview. We tried catching up with Dario to tell us something about himself. He usually spends his day deep in data, but he’s also running the business development team and you can often find him jumping from one meeting to another. So, when we managed to catch him, we asked him to tell us something more about his work and his daily activities. Tell us who you are and what you do in Digital Poirots? I am Dario, a data scientist, data engineer, and data team lead. What does your typical day look like? Well, I don’t have a typical day. 🙂 Sometimes I spend more time on documents and business-related stuff (HR, marketing,…), but I try to invest as much time in team and technology.  When working on a project for a client my day starts with a quick status overview, and a sync with team members after which the actual work starts. The work is hard to explain in a few lines, as it can be everything from designing data pipelines, and data warehouses, or implementing machine learning or statistical models. In general, you inspect a lot of data sources (files, APIs, database tables), process them, combine them, calculate new metrics based on business rules, or use them for predictions. Internal projects serve us to learn and upgrade our knowledge but require planning and management. That’s where I step in. There is a lot of client communication involved in everyday work, especially when working on data projects with complex business logic in the background. We are on a constant lookout for talent so part of my day (not every day) is preparing for interviews and interviewing them. What drives you in your work? There are several factors. I remember one situation when I started working as a data engineer and solved some issues in the database that unblocked the whole web development team. I felt great as my work had an effect in the real world, it is so much different than working on a university project that gets discarded the next minute after the submission.  Also, client satisfaction after consolidating their datasets, or visualizing data and insights in dashboards is important to me, as it reflects their everyday business greatly. What’s the best part about your job? To me, the best part is that the job is not repetitive. Principles stay the same, but data and the application of data changes. What don’t you like about your work?  I dislike days with too many meetings. Which project you’ve worked on is your favorite? Regarding data engineering, definitely My Dairy Dashboard project. For data science, I pick Milk Forecasting and Bellabeat projects. What’s the biggest mistake you’ve made in your career or what ups moment you had? The first day after a vacation in 2020 I messed up tables in the data warehouse in the test environment and it required a full historical load and some manual work to fix it. What drives you crazy about your job or in your daily activities? Probably multitasking too many totally different and opposite things. Which technology/tech stack do you like the most? SQL, Python and its ML stack (Pandas, Scikit-learn), and Apache Spark. What advice would you give to someone entering this field? Don’t focus too much on a specific technology, but learn concepts and get experience by working on real-world projects. Technology is just one aspect of everyday work in the data field. Get experience in presenting your data ideas to non-technical people. When things are unclear, propose a few solutions and present them to stakeholders, instead of just asking them what to do. Your favorite person to work with? Me, myself, and I! Haha The most time I spend working with my brother Renato. If you can compare your job to one movie or show, what would it be? Most of us are fans of the show “The Office” so we like to joke around that we’re like that, especially when we say bad puns or jokes. If you can choose one song to go along with your job or which makes you get in the zone while working, what would it be? When I am in the zone I prefer light instrumental music. Back in the day, we had Radio Sljeme playing in the office on low volume. Stay tuned for other interviews to meet all our other data detectives!

How to turn SKU analysis to your advantage

SKU analysis

SKU or Stock Keeping Unit is a base code for tracking goods or products across various functions or departments inside or outside the company. Based on the unique identification number, businesses follow the movements of specific products or product categories to better understand their performance and profitability. When companies deal with a great number of different products, SKUs allow them to make a distinct difference between them to minimize errors or duplicate data. From a production point of view, each finished product gets an assigned SKU entered into the system with descriptions. This code is used to move goods in the system from production to warehouse and later on to sales and distribution. From a sales point of view, SKUs are used to differentiate assortments and assign specific codes according to different product attributes. It is used in sales efforts, marketing, catalogs, and inventory management.  Often, when dealing with a bigger assortment, it’s hard to specify which products are more cost-effective and necessary to keep in supply. That requires daily analyses across multiple databases with filtering out the only necessary information. SKUs allow for analysis across individual products, brands, or product categories, depending on each department’s needs and requirements. For each department, there is a different way of interpreting SKU’s performance. Certain metrics require certain data sets, while others require different ones. There are vast methods of utilizing SKU data and they all depend on set goals. What does SKU do If we track SKU from production, data that accompanies it can be information about logistics (size, packing, other product characteristics, and attributes), cost, BOM (bill of materials), production cycle length, and quality control. In warehouse management, SKUs have data on pallets, packing size, quantity, inventory, warehouse pallet address, usage or expiration data (if it’s perishable products), and time of warehouse entry or exit. In sales, data that is tracked is sales numbers, price, quantity, date, customer, location, and cost of sales efforts. For marketing, SKUs are recorded through sales numbers, marketing cost, and marketing and trade marketing data (information about product characteristics, brand, country of origin, etc.).  Each piece of information is directly linked to the SKU. Often the information about SKUs is kept in Master Data Management software, specifically in the Product Information Management part of the software to track static data. Dynamic data can be tracked also through MDM or ERP systems. With so much data that goes along with SKUs, its importance can, unfortunately, be disregarded or even not fully utilized. Some metrics and insights can be drawn from it to explain business operations and spot errors or opportunities. One of the most important data is sales numbers that explain the effectiveness and profitability of each SKU or product. They are calculated based on orders and goods delivered, or from revenue. Price and sold quantity are a basis for revenue or earnings calculation and can explain the demand for some products. If compared to dates, they can show trends across different time periods. SKU sales analysis To start, let’s delve into the wholesale or retail part of SKU analysis. It’s a specially interesting part of the SKU movement that is directly related to the supply chain. Here, SKU movements are directly linked to sales and inventory management. Data extracted there is price, sold quantity, location, product category, and date. With that information, you can do a full analysis of the assortment and its profitability while interpreting with consideration of geographical and time factors. From sold quantities and prices, one can determine the sales for that time period. If connected to specific dates, it can point out outliers and alert users that some SKUs are showing unordinary over or underperformance when compared to previous time periods. When you look at one SKU you can analyze sales trends movement over time and compare it to other SKUs or product categories in the same or different locations.  This approach helps determine which products should be kept in stock and pushed more in sales. If according to trends, can sellers expect an increase in demand, they automatically know when to increase stock to satisfy the needs of customers. It’s also a major part of demand planning. Demand planning is based on the number of sold units over time. With historical data in place, it can forecast future trends in demand for certain SKU or product categories. When analyzing SKU performance through time, it can be determined if the product is new on the market so a phase-in is needed, or if it’s in the last stage of the product life cycle and should be, therefore, put in the phase-out process. With demand planning, inventory management could be more optimized and cost-effective. The same principle can be applied to determine products in the assortment that are below the category performing line or generally have low sales. It could be an indicator to retract the product from stores or to intensify marketing and sales strategies. If there are low sales of the complete category, it could be an incentive to pull it off the shelves completely to lower costs and avoid overstocking.  SKU analysis could also be used in price management and optimization. When comparing SKU prices to other SKUs or category averages, sometimes there is room for price increase without losing out on sales. But the main part is analyzing if the number of sold units is closely correlated to the price change or not. This helps determine if consumers are price-sensitive when talking about that product or product category. From a market share perspective, SKU analysis is beneficial when trying to determine which products or product categories make up the most share in sales. Certain locations are going to buy more of some SKUs while others are going to do the same for some other products. This is a signal for sellers to plan their store assortment and keep more goods that are higher in demand while reducing the stock for those with low or zero demand. Why focus on

Data, but make it customer-centric

Customer centric

3 methods that lead to your perfect customer Looking at any data set you collect in your business processes, you’ll realize it revolves around your customer or client. Each piece of information provides insights that are used for the objective of providing the best user experience. Whether we talk about products or services, the goal is the same – to sell them. But there has been a shift a long time ago from product-centric business models to customer-centric. If you ask yourself why – the answer lies in data showing how consumers interact with companies. It isn’t anymore about just products or services doing their job, it’s the value and experience consumers get from using them. And to get to that, everything starts with them – your customers. Companies have invested so much time and money in understanding consumer behavior and who uses their products and services and how. Basic principles of segmentation have been adapted to discover who brings the most value to the business. But there are also methods for discovering who doesn’t. Let’s look at one of the industries where we had the challenge of delving deep into data to recognize customer segments by certain attributes and standards. The insurance industry is a highly competitive market, from car insurance, and property to health, the battle for customers is constant. It’s an area where price and service value play a great role. But it’s not only client acquisition that poses a challenge. It is how you determine which customers or clients are valuable to you. So, let’s dive into a couple of areas where data helps in discovering who the clients are in the insurance industry, how they behave, and how you can determine what level of value they bring. Insurance users segmentation From internal data sources, CMRs, sales leads and records, and such, insurance companies have access to data on their past and current customers. If we talk about the B2B segment, which we did in the challenge, the size, profile, and revenue of the company (firmographic attributes) are factors that could divide them into appropriate groups. It’s about understanding clients, which insurance services or products they use, and what’s important to them when they choose them. Even insurance packages or deals are made tailored to clients’ needs and wants, and according to their shared attributes. But it all starts from data that is centered around them. The process starts with data cleansing and analysis to determine what will influence the segmentation and provide satisfactory results.  Using data science methods and procedures provided us with the means to do the segmentation effectively and more precisely. In the end, this allows the insurance company to target its clients more effectively, and to create more cohesive and straightforward sales, marketing, and product strategies. There is less unnecessary resource expenditure on processes that are now covered by data-driven solutions. User behavior forecasting Clients and consumers are not static. Their behavior and attributes change over time, and so does their willingness or need to invest in certain products and services. By completing the first step, market segmentation, the insurance company now has the basis to track future market movements, either by segment or attributes.  Once a golden customer, could prove to be the exact opposite in the future. His firmographic characteristics could change. Revenue could increase or decrease, the size of the company may change, and it all affects the usage of insurance and the price itself. By implementing forecasting solutions, we can predict how a certain segment will behave based on historical data. We can determine if the segment size will change and where are better opportunities to offer products or services. It can help in predicting which segments could make insurance companies lose money if they don’t shift the focus to profitable ones. Predicting user behavior will allow any company for that matter to be more cost-effective and recognize the needs of the market on time. Client pruning Client pruning is closely connected to consumer data and segmentation. How do you determine which clients provide the most value and which don’t if you never delve deep into their data? Pruning is a technique reserved for recognizing which current clients or customers are profitable and will be profitable in the future and which ones will only invoke costs and lost opportunities. Client pruning helps management answer the question of who they want to serve. It’s used to maximize profitability and center efforts on certain segments.  It starts with collecting and analyzing data connected to client quality and profile. Types of data used could be the same or similar to the segmentation ones, like company size, revenue, and the number of employees, which is important for insurance services. But grading clients could be done by analyzing factors such as pricing or total client size, timely payment, number of insurance services they use, profitability, last point of contact, number of sales reach-outs and success rates, cross-selling opportunities, and referral capabilities, and internal notes on client behavior. With data science capabilities, clients can be then ranked from those that are profitable or show potential to be profitable to those that mainly generate costs. Implementing a forecasting feature could help in anticipating which client will need to undergo pruning. Customer-centric is data-centric Data is a powerful tool. And only by using data science and data engineering methods do we see and utilize its full potential. The market is and will continue to be customer-centric because companies don’t sell products or services, they sell experiences. So, how do you reach that? By undergoing these three steps, segmentation, forecasting, and pruning. By doing so, companies can focus on important customers and adapt their strategies to serve only them. It will create a joint point of customer objectives and company objectives, where each side has its needs met. Such equilibrium will provide for a stable and more effective collaboration and relationship.

The endless possibilities of machine learning

Machine learning

Machine learning – such a powerful and popular term that has taken over the software development industry, especially data-driven software solutions. It’s a method of data analysis that is based on the concept that systems can learn from data, identify patterns, and forecast future movements. For many companies that either generate or collect vast amounts of data, machine learning can provide a window into the future on how will certain parts of the market behave.  The applicability of machine learning can range from simple predictions to more complex ones. It can be used in different industries and markets, depending on the objective. Our experience with it is that even though getting precise predictions can be hard and sometimes daunting, it also allows for more possibilities to maximize data utilization. Machine learning applications in the manufacturing industry Even though our use cases have been used in farming, their concept and machine learning methods can be applied across various industries. No matter the commodity or product, the basic principle stays the same, and what we know and what we’ve learned can only help in making the next model more accurate and better performing.  Two of the main solutions where machine learning was applied are input and output forecasting and those solutions are adaptable to any industry and production. Those two solutions predict future outcomes on input or output quantity, quality, and results. But, what are some cases where your output is dependent on changeable and fluctuating factors? In farming, you as a producer are counting on your animals’ best performance and you have to be on top of the game to keep those animals in good condition.  In the commodities industry, like milk, cheese, and other dairy products, you are heavily influenced by cows and their state on the farm. The quantity and quality of dairy products, mainly milk, are the result of the cow’s life cycle and productive phase. A producer’s output is defined by the number of animals on that farm and how many of them can provide satisfactory results. To make the future a bit more stable and adaptable to farm and market changes, some of the cases regarding machine learning were used to predict how animals will behave. It was important to have a model that could give an overview of animals’ performance. Not many are familiar with each factor that could influence milk production and cows. That’s why our machine learning models were focused on doing just that – predicting factors that are a part of a cow’s productivity.  Can you predict declines in animal products? Our first model was applied in the breeding analysis. In times of breeding, a cow cannot provide milk that will be used in future production. So, it was important to determine when a certain cow will not be available, so farmers can adjust expectations and planning on milk quantities. Based on the historical data of past breeding and breeding processes, it can be predicted when the animal will not give milk. It also helps predict when a cow will be in the breeding phase, so farmers can anticipate and plan for the birth of calves. They can also plan when is the best time for breeding if the goal is to grow the number of animals on the farm. The next model covers potential diseases in cows. If an animal is sick or has some health issues, it will not be able to provide results. Predicting such occurrences is beneficial for farmers so they can react on time. For example, mastitis occurrence in cows will affect milk production. If farmers can predict future occurrences on time, they can react quickly regarding cow’s health and milk quality. Based on historical data on disease and data on cow’s life cycle, breed, and such can help build models to anticipate mastitis.  Machine learning can also be applied to track an animal’s life cycle. It can predict survival by the season of birth or even anticipate the length of its life. If you track data on all influences on animals and their personal information on health, breed, or productivity, you can build models that can forecast when a certain animal’s life or health is in decline.  Machine learning applicability can be seen through the whole life cycle of an animal. You can use data collected through their life, to help you understand how it will behave in the future and what is their productivity, meaning how much product, or in our case milk, they can provide. Why machine learning? Such models help minimize uncertainty and allow for quicker reactions that will help stabilize and optimize production. They provide insights and metrics that give an overview of the whole production and how will it operate in the future. Machine learning helps create scenarios where users can spot potential problems, errors, and disruptions. It’s a tool whose goal is to help prepare the user for the future. The usage of ML is on the rise and it will continue so. The possibilities it offers are endless, and it can utilize data to your biggest advantage. In a world that revolves around information, the next step is to use it to predict future outcomes. Any company that devels into such endeavors will certainly benefit in long term, especially in establishing its competitive advantage.