November 26, 2024
When Qlik was founded in 1993, hard drives were measured in megabytes, and the Internet was primarily text-based. If lucky, you could get information in structured columns and formats.
Fast forward thirty years, and some estimate YouTube alone has 4.3 petabytes of data loaded every day.
The federal government certainly has its share of formatted data. A recent survey showed that 80% of data collected by the federal government is unstructured. This is information like text files, videos, or emails that are stored in many formats. As a result, it isn’t easy to store and manage.
This has a real impact when an organization tries to take advantage of Artificial Intelligence.
"AI is a long game. You know, the majority of the AI initiatives are very exploratory right now in the federal government, in public sector in general. And you know, we're, we're seeing a huge amount of interest in, in one profiling data, saying what do I have governing the data, and classifying it. "
Andrew Churchill, Qlik Tweet
Today, we sit down with Andrew Churchill to discuss creating a solid data foundation for AI. We detail topics like data movement, data streaming, and data quality during the discussion.
He differentiates between data lakes and data warehouses as strategies for handling all the unstructured data used for training AI models.
If you enjoyed this episode, then you may want to listen to
EP-205: AI = Data + Compute + Algorithm
Reinforces the technical components discussed in EP-198, especially the importance of aligning data and computing infrastructure before deploying AI.
EP-202: Innovation in Handling Complexities of AI
Explores how federal agencies can manage AI’s complexity, a great follow-up to foundational strategies.
