Skip to content

Details

Data Science is an exploding field across a range of industries. The general paradigm is for data scientists to work on source data captured in Data Lakes/Warehouses/Vaults/Marts; or, less often, accessing the raw data at source.
The industrial data sources usually originate from OLTP (Online Transaction Processing) applications, such as ERP/Financials: or Semi-structured/Unstructured data sources, such as Twitter feeds.
Using dashboards/reports or AI/ML, data scientists have to analyse this vast body of source data to glean Business Intelligence (BI) and analytics. We pose the question: are underlying data models required on the data held within Data Lakes in order to get meaningful BI analytics?
Data Lakes generally implement the Delta Lake medallion architecture consisting of three layers, as follows:
1. Bronze: This is where raw data from all sources are stored as created without any manipulation of any kind
2. Silver: This is a curated layer on the raw data which validates/de-duplicates the data from the Bronze layer
3. Gold: This is business consumption layer that restructures /aggregates the Silver layer to create artefacts ready for end-users, including data scientists.
AI/ML can be used to profile the source data, especially if the volumes are very large. For instance, data profiling can discover whether a Foreign-Key (FK) relationship exists between the Customer and Account tables; and whether the relationship is mandatory or optional. This would be incorporated in the data model.
We therefore maintain that in order for a data scientist to undertake BI analytics, it is necessary to build a relevant, underlying data model. This allows data from various sources to be integrated and targeted metrics computed. Without a data model, there is no meaningful way to relate the data held in the Data Lake.

Hybrid Session

As well as getting back in person, we have a Teams link available here. Please answer the question when you register about whether you’ll be there in person or remotely.

What Else to Expect?

The chance to network. We plan to incorporate the best parts of SQL Social and the SQL Server Group in our User Group, so there will be networking time, learning opportunities and introductions to new ideas.

Pizza and Door Prizes

Yes, we’d love to see you back in person if you are comfortable doing so. As a reward, we’ll have pizza and door prizes for those who come along and join in the networking time.

Agenda

5.30 Networking and Pizza
6.00 Why we need Data Modelling in Data Science?

Related topics

Events in Southbank
Microsoft Azure
SQL Azure
Data Analytics
Data Visualization
Business Intelligence & Data Warehousing

You may also like