top of page
Search

Practical Machine Learning & AI

Oct 20, 2021
5 min read

Last fall when I took on ShippingEasy’s machine learning problem, I had no practical experience in machine learning. Getting such a task put on my plate was somewhat terrifying, and even more so as we started to wade into the waters of machine learning. Ultimately, we overcame those obstacles and delivered a solution that allowed us to automate our customer’s actions with greater than 95% accuracy. Here are some of the challenges that we experienced when applying machine learning to the shipping & fulfilment domain, and how we broke through them.


Lost in Translation

Machine learning is a subfield of computer science stemming from research into artificial intelligence.[3] It has strong ties to statistics and mathematical optimization, which deliver methods, theory and application domains to the field.

So sayeth the Wikipedia. These roots are where the lexicon of the machine learning stems from. If you have not been working directly in machine learning, statistics, math or AI, or perhaps your exposure to these are long past, a discussion about machine learning will be hard to follow. Often it is taken for granted that you know what classification, regression, clustering, supervised, unsupervised and a host of other terms mean.


As a result, you will be somewhat lost until you can get familiar with this language. Getting a good book will help. I would recommend Machine Learning, a short course. It clocks in at less than 200 pages, and so is something that a working professional can consume. Even if you can’t follow everything in the book, reading through it will give you a foundation that will allow you to make use of all the other resources you may find online.


To give you a starting point to building your vocabulary, I will offer a few terms here that will help determine what type of machine learning problem you are dealing with.


Supervised vs Unsupervised Learning: Supervised learning is where you have a set of input data with known outcomes by which you wish to predict the outcome of future inputs. Our problem at ShippingEasy was of this type. We had past orders and shipments and needed to predict shipments given future orders. Unsupervised learning is where you have input data, but no known outcomes. You are searching for what features have meaning within a set of data. If graphed, the data will form clusters around the patterns of meaningful features.


Classification vs Regression: Within supervised learning, there are problems of classification and regression. Classification is where you wish to determine the class (output) of an input. For instance, predicting what shirt color a person may wear on a given day based on data about what shirts they have worn in the past. The different shirt colors are the classes that you are attempting to predict.


Regression is where you wish to determine a numeric value given other numeric inputs describing the sample. For instance, predicting an engineer’s salary based on age and years in the industry. Given enough past data, you could given a new age and years in the industry arrive at a salary that would statistically be close to correct (assuming true relationships between age, years in industry and salary).


Algorithmic Obsession

Once you have a foundation of concepts and language, you can start looking into all of the amazing resources on the web for machine learning. Stanford’s machine learning videos are great, as are mathematicalmonk’s youtube videos.


These are fantastic resources for learning how to write machine learning algorithms. But these turned out to not be of much use to me. Not because they are not great, but because the practical application of machine learning is about solving a domain problem, not writing machine learning algorithms. To make the point, consider this portion of a machine learning algorithm expressed in mathematical notation (which is yet another barrier to the uninitiated):


What does this have to do with the problem you are trying to solve? Absolutely nothing. This is a description of one portion of an algorithm that may be fed arbitrary data to produce statistically relevant results. It has been implemented by someone smarter than you in an open source library or possibly a service offering. You could implement it perfectly, and it could produce great results or really bad results. It all depends on the relevancy of the data you feed to it, which brings me to my last point.


Its the Data, Stupid

While there are a tremendous number of resources for how to write machine learning algorithms, there are not many dealing with how to find relevant data within a domain that will allow an algorithm to produce accurate results. This is where you will find that you have spent most of your time, effort and creativity at the end of an applied machine learning project if you were smart enough to use a good machine learning library or service.


That algorithms dominate the resources for machine learning makes a certain amount of sense. Algorithms are generic and have practicability for many different scenarios. The K-Nearest Neighbor algorithm may be able to predict what movies you would like to watch on Netflix, or it might be able to predict which sex offenders are at high risk for recidivism. These different applications of K-Nearest neighbor would have very different data that needs to be surfaced from their respective domains and fed to them, however.


There exists an area of machine learning geared towards feature detection, and I won’t dismiss its validity. I will say, however, that if someone understands the domains of movie consumption and purchasing dynamics or criminal behavior, justice and rehabilitation, they have a leg up in practically applying machine learning to those domains. For even if there is a statistical correlation between day of the week and movie choices, it does not mean that there is a causative relationship between them.


Some of the data will be obvious. It winds up being a value in a column of a row in the database and it screams its pertinence. Some will be much less obvious and need to be inferred. For instance, for the sex offender recidivism problem, there are probably a number of criminal incidents, each with a timestamp for when they occurred. For any given person, the amount of time that has passed since their last criminal event, in days, might need to be calculated and included with the data sent to the algorithm. This ‘freshness’ of their criminal activity needs to be inferred from your data, and it may be a key to getting the desired results in predicting future likelihood of behavior. Or it might not.


I think the moral of the story here is that to really apply machine learning in a practical way, being a mathematical or statistical wizard is not the most important element of success. What I feel is more important is having an understanding of the domain to know what data is relevant and an explorer’s curiosity to have meaningful hunches and a willingness to explore and vet them. You will need to be comfortable employing something resembling a scientific method – ensuring accuracy is measurable, quantifying the effects of change, and meticulously exploring isolated changes to discover what data affects a system.


In conclusion

Employing machine learning to solve domain problems can provide huge value to a company or the public at large. Learning machine learning and how to properly apply it to a domain, however, can be challenging. You will need to develop a knowledge of the fundamentals of machine learning, but do not need to be a compsci, math or statistics guru to employ it. Leverage existing libs or services, and then focus your efforts on finding the meaningful data within the domain, both obvious and obscure, that will allow satisfactory results to be achieved.


 
 
 

10 Comments


b9tmuha2hukj5cr
3 hours ago

It is fascinating to see how quickly sztuczna inteligencja is moving from experimental technology into tools that people can actually use in their everyday work. The most valuable applications seem to be those that save time without making the process unnecessarily complicated. I expect education and business will both change considerably because of this technology.

Like

b9tmugzo8f2exqm
8 hours ago

The way companies approach reklama has changed significantly in recent years, especially as audiences spend more time across different digital platforms. What I find most interesting is that good advertising still depends on understanding the customer rather than simply increasing the budget. A useful overview of a topic that continues to evolve quickly.

Like

Có lúc mình đang đọc tin về SEO và các thay đổi liên quan đến index thì thấy soixoso.net xuất hiện trong danh sách mình đang xem. Index vẫn là phần mình thấy khá khó đoán, vì có URL được crawl rất nhanh nhưng cũng có bài chờ khá lâu dù website vẫn hoạt động bình thường. Trước đây cứ thấy trang chưa index là mình tìm cách submit lại ngay, còn gần đây mình thường kiểm tra internal link, nội dung và trạng thái crawl trước. Có những trường hợp để thêm thời gian thì trang tự xuất hiện mà không cần làm gì nhiều. Vì thế mình đang cố phân biệt vấn đề kỹ thuật thực sự với những…

Like

Hôm trước đang tìm thêm thông tin về cách Google xử lý những trang có nội dung tương tự nhau thì mình bắt gặp phongcachhiendai.net. Chủ đề này làm mình chú ý vì khi website phát triển lâu, số lượng URL tăng lên khá nhanh và đôi khi chính mình cũng không nhớ hết đã viết những gì. Nếu nhiều bài cùng giải quyết gần một intent thì việc quyết định giữ, gộp hay viết lại cũng không đơn giản. Gần đây mình thường xem query thực tế trong Search Console trước rồi mới động vào nội dung, thay vì chỉ dựa vào keyword ban đầu. Cách này giúp nhìn rõ hơn Google đang hiểu từng URL theo hướng nào.…

Like

Mình tình cờ gặp echoreach.net trong lúc đang xem một số tin tức và thảo luận mới về SEO. Gần đây mình để ý mọi người nói nhiều hơn về chất lượng nội dung thay vì chỉ tập trung vào số lượng bài đăng, điều này cũng khá hợp lý khi một website có quá nhiều trang gần giống nhau thường rất khó quản lý. Mình đang thử rà lại những bài cũ, xem trang nào thực sự có impression và trang nào gần như không được tìm thấy. Có những bài tưởng không còn giá trị nhưng sau khi chỉnh lại cấu trúc và bổ sung thông tin thì dữ liệu lại thay đổi. Mình chưa thử trên đủ nhiều…

Like
Post: Blog2 Post

7193717576

©2021 by devquixote. Proudly created with Wix.com

bottom of page