Skip to main content

From Unstructure Data to Data Model

Collecting and preparing unstructured data for data modelling involves several steps. Here's a step-by-step guide with a basic example for illustration:

Step 1: Define Data Sources

Identify the sources from which you want to collect unstructured data. These sources can include text documents, images, audio files, social media feeds, and more. For this example, let's consider collecting text data from social media posts.

Step 2: Data Collection

To collect unstructured text data from social media, you can use APIs provided by platforms like Twitter, Facebook, or Instagram. For this example, we'll use the Tweepy library to collect tweets from Twitter.


import tweepy

# Authenticate with Twitter API

consumer_key = 'your_consumer_key'

consumer_secret = 'your_consumer_secret'

access_token = 'your_access_token'

access_token_secret = 'your_access_token_secret'

auth = tweepy.OAuthHandler(consumer_key, consumer_secret)

auth.set_access_token(access_token, access_token_secret)

# Initialize Tweepy API

api = tweepy.API(auth)

# Collect tweets

tweets = []

usernames = ['user1', 'user2']  # Add usernames to collect tweets from

for username in usernames:

    user_tweets = api.user_timeline(screen_name=username, count=100, tweet_mode="extended")

    for tweet in user_tweets:


# Now, 'tweets' contains unstructured text data from social media.


Step 3: Data Preprocessing

Unstructured data often requires preprocessing to make it suitable for modeling. Common preprocessing steps include:

- Tokenization: Splitting text into individual words or tokens.

- Removing special characters, URLs, and numbers.

- Lowercasing all text to ensure uniformity.

- Removing stop words (common words like "the," "and," "is").

- Lemmatization or stemming to reduce words to their base forms.

Here's an example of data preprocessing in Python using the NLTK library:


import nltk

from nltk.corpus import stopwords

from nltk.tokenize import word_tokenize

from nltk.stem import WordNetLemmatizer'punkt')'stopwords')'wordnet')

# Example text

text = "This is an example sentence. It contains some words."

# Tokenization

tokens = word_tokenize(text)

# Removing punctuation and converting to lowercase

tokens = [word.lower() for word in tokens if word.isalpha()]

# Removing stopwords

stop_words = set(stopwords.words('english'))

filtered_tokens = [word for word in tokens if word not in stop_words]

# Lemmatization

lemmatizer = WordNetLemmatizer()

lemmatized_tokens = [lemmatizer.lemmatize(word) for word in filtered_tokens]

# Now, 'lemmatized_tokens' contains preprocessed text data.


Step 4: Data Representation

To use unstructured data for modeling, you need to convert it into a structured format. For text data, you can represent it using techniques like Bag of Words (BoW) or TF-IDF (Term Frequency-Inverse Document Frequency).

Here's an example using TF-IDF representation with scikit-learn:


from sklearn.feature_extraction.text import TfidfVectorizer

# Example list of preprocessed text data

documents = ["this is an example document", "another document for illustration", "text data preprocessing"]

# Create a TF-IDF vectorizer

tfidf_vectorizer = TfidfVectorizer()

# Fit and transform the text data

tfidf_matrix = tfidf_vectorizer.fit_transform(documents)

# Now, 'tfidf_matrix' contains the TF-IDF representation of the text data.


With these steps, you've collected unstructured data (tweets), preprocessed it, and represented it in a structured format (TF-IDF matrix). This prepared data can now be used for various machine learning or data modeling tasks, such as sentiment analysis, topic modeling, or classification. Remember that the specific steps and libraries you use may vary depending on your data and modeling goals.

Photo by Field Engineer


Popular posts from this blog

Financial Engineering

Financial Engineering: Key Concepts Financial engineering is a multidisciplinary field that combines financial theory, mathematics, and computer science to design and develop innovative financial products and solutions. Here's an in-depth look at the key concepts you mentioned: 1. Statistical Analysis Statistical analysis is a crucial component of financial engineering. It involves using statistical techniques to analyze and interpret financial data, such as: Hypothesis testing : to validate assumptions about financial data Regression analysis : to model relationships between variables Time series analysis : to forecast future values based on historical data Probability distributions : to model and analyze risk Statistical analysis helps financial engineers to identify trends, patterns, and correlations in financial data, which informs decision-making and risk management. 2. Machine Learning Machine learning is a subset of artificial intelligence that involves training algorithms t...

Wholesale Customer Solution with Magento Commerce

The client want to have a shop where regular customers to be able to see products with their retail price, while Wholesale partners to see the prices with ? discount. The extra condition: retail and wholesale prices hasn’t mathematical dependency. So, a product could be $100 for retail and $50 for whole sale and another one could be $60 retail and $50 wholesale. And of course retail users should not be able to see wholesale prices at all. Basically, I will explain what I did step-by-step, but in order to understand what I mean, you should be familiar with the basics of Magento. 1. Creating two magento websites, stores and views (Magento meaning of website of course) It’s done from from System->Manage Stores. The result is: Website | Store | View ———————————————— Retail->Retail->Default Wholesale->Wholesale->Default Both sites using the same category/product tree 2. Setting the price scope in System->Configuration->Catalog->Catalog->Price set drop-down to...

How to Prepare for AI Driven Career

  Introduction We are all living in our "ChatGPT moment" now. It happened when I asked ChatGPT to plan a 10-day holiday in rural India. Within seconds, I had a detailed list of activities and places to explore. The speed and usefulness of the response left me stunned, and I realized instantly that life would never be the same again. ChatGPT felt like a bombshell—years of hype about Artificial Intelligence had finally materialized into something tangible and accessible. Suddenly, AI wasn’t just theoretical; it was writing limericks, crafting decent marketing content, and even generating code. The world is still adjusting to this rapid shift. We’re in the middle of a technological revolution—one so fast and transformative that it’s hard to fully comprehend. This revolution brings both exciting opportunities and inevitable challenges. On the one hand, AI is enabling remarkable breakthroughs. It can detect anomalies in MRI scans that even seasoned doctors might miss. It can trans...