An Overview of Building a Merchant Name Cleaning Engine with SequenceMatcher and spaCy

https://cdn-images-1.medium.com/max/2600/0*VxuNEHPDKJXjhY6L

Original Source Here

Layer 3 : Merchant Names Cleaning with spaCy

By completing first two layers, we are able to solve some of the merchant names cleaning problems such as names of mis-spelling, different cases, missing characters/spaces and even some of the non-messy merchant inputs by simply returning a similarity score table.

However, we are actually still in the phase of working with a rule-based cleaning engine, which means so far we still haven’t learned from the data. Furthermore, even by use of a typical machine learning model, the training phase still requires a large amount of time to perform feature engineering in creating more informative features.

… potentially informative transaction-level features such as dollar amount and category, while also generating word-level natural language features such as word position within the label (e.g., 1st, 2nd), word length, proportion of vowels, consonants, and alphanumeric characters, among others.

CleanMachine: Financial transaction label translation for wallet.AI

Therefore, I researched on how to use a deep learning model to create the cleaning engine. The advantage of using a deep learning model is that we are able to “skip” the feature engineering step and let the model itself detect any insightful patterns from the inputs.

Introduction to spaCy

A free short course on spaCy can be found as following:

According to the spaCy Guide:

spaCy is a library for advanced Natural Language Processing in Python and Cython. It’s built on the very latest research, and was designed from day one to be used in real products. spaCy comes with pre-trained statistical models and word vectors, and currently supports tokenization for 60+ languages.

It features state-of-the-art speed, convolutional neural network models for tagging, parsing and named entity recognition and easy deep learning integration. It’s commercial open-source software, released under the MIT license.

Since the merchant names cleaning problem can be classified under the topic of named entity recognition(NER), I feel confident that a spaCy model will have a good performance by feeding in a set of representative input data.

Training a spaCy’s Statistical Model

https://spacy.io/usage/training#section-basics

To train a spaCy model, we don’t just want it to memorize our examples — we want it to come up with a theory that can be generalized across other examples.

Therefore, the training data should always be representative of the data we want to process. For our project, we may want to select training data from different types of merchant names. Eventually, our training data will be in form of an entity list like the following:

TRAIN_DATA = 
[
('Amazon co ca', {'entities': [(0, 6, 'BRD')]}),
('AMZNMKTPLACE AMAZON CO', {'entities': [(13, 19, 'BRD')]}),
('APPLE COM BILL', {'entities': [(0, 5, 'BRD')]}),
('BOOKING COM New York City', {'entities': [(0, 7, 'BRD')]}),
('STARBUCKS Vancouver', {'entities': [(0, 9, 'BRD')]}),
('Uber BV', {'entities': [(0, 4, 'BRD')]}),
('Hotel on Booking com Toronto', {'entities': [(9, 16, 'BRD')]}),
('UBER com', {'entities': [(0, 4, 'BRD')]}),
('Netflix com', {'entities': [(0, 7, 'BRD')]})]
]

The training data I choose is just a sample. The model can take more complex inputs. However, it can be a little boring to annotate a long list of merchant names. I would like to recommend another data labeling tool UBIAI to complete this task, as it supports output in a spaCy format or even in an Amazon Comprehend format.

It may require some experience on how to select a representative data input. As you practice more and observe the way how a spaCy model learns, it will become clearer that “representative” probably means “different locations”. It’s the reason why we need to provide an entity start & end index in the input data, because it can help the model to learn the patterns from different contexts.

If a model is often trained with the location of the first word being a merchant name (Amazon ca), then it tends to believe that a merchant name only locates at the beginning of an input. This will likely cause a bias and result in a wrong prediction for input such as “Music Spotify” because “Spotify” happens to be the second word.

However, it’s also important to include various merchant names in the input. Just be aware that we don’t want our model to merely memorize them.

Once you have finalized tuning your training data, the remaining process will almost be automated.

import spacy
import random
def train_spacy(data,iterations):
TRAIN_DATA = data
nlp = spacy.blank('en') # create blank Language class
# create the built-in pipeline components and add them to the pipeline
# nlp.create_pipe works for built-ins that are registered with spaCy
if 'ner' not in nlp.pipe_names:
ner = nlp.create_pipe('ner')
nlp.add_pipe(ner, last=True)
# add labels
for _, annotations in TRAIN_DATA:
for ent in annotations.get('entities'):
ner.add_label(ent[2])
# get names of other pipes to disable them during training
other_pipes = [pipe for pipe in nlp.pipe_names if pipe != 'ner']
with nlp.disable_pipes(*other_pipes): # only train NER
optimizer = nlp.begin_training()
for itn in range(iterations):
print("Statring iteration " + str(itn))
random.shuffle(TRAIN_DATA)
losses = {}
for text, annotations in TRAIN_DATA:
nlp.update(
[text], # batch of texts
[annotations], # batch of annotations
drop=0.2, # dropout - make it harder to memorise data
sgd=optimizer, # callable to update weights
losses=losses)
print(losses)
return nlp
prdnlp = train_spacy(TRAIN_DATA, 20)# Save our trained Model
modelfile = input("Enter your Model Name: ")
prdnlp.to_disk(modelfile)
#Test your text
test_text = input("Enter your testing text: ")
doc = prdnlp(test_text)
for ent in doc.ents:
print(ent.text, ent.start_char, ent.end_char, ent.label_)

Above code is from the following Medium article, as I found it very helpful and it inspired me to test spaCy on a merchant names cleaning problem.

Evaluating a spaCy model

By successfully completing the training step, we can monitor the model progress by checking its loss values.

Statring iteration 0
{'ner': 18.696674078702927}
Statring iteration 1
{'ner': 10.93641816265881}
Statring iteration 2
{'ner': 7.63046314753592}
Statring iteration 3
{'ner': 1.8599222962139454}
Statring iteration 4
{'ner': 0.29048295595632395}
Statring iteration 5
{'ner': 0.0009769084971516626}

The model is then shown the unlabelled text and will make a prediction. Because we know the correct answer, we can give the model feedback on its prediction in the form of an error gradient of the loss function that calculates the difference between the training example and the expected output. The greater the difference, the more significant the gradient and the updates to our model.

To test your model, you can run the code cell below:

#Test your text
test_text = input("Enter your testing text: ")
doc = prdnlp(test_text)
for ent in doc.ents:
print(ent.text, ent.start_char, ent.end_char, ent.label_)

For instance, we can use “paypal payment” as input and test the model if it can detect “paypal” as the correct brand name.

The model did a great job considering PayPal did not appear in the training input.

This also concludes my project of building a merchant names cleaning engine with a spaCy model.

AI/ML

Trending AI/ML Article Identified & Digested via Granola by Ramsey Elbasheer; a Machine-Driven RSS Bot



via WordPress https://ramseyelbasheer.wordpress.com/2021/01/31/an-overview-of-building-a-merchant-name-cleaning-engine-with-sequencematcher-and-spacy/

Popular posts from this blog

Fully Explained DBScan Clustering Algorithm with Python

The 2021 machine learning, AI, and data landscape

Hierarchical clustering explained