---
title: "Resources | Shipshape Data"
description: "Practical guides on data foundations, governance and AI delivery from Shipshape Data. Written by the team that builds these systems for clients."
canonical: https://shipshapedata.com/resources/
language: en-GB
---

# Guides for getting your data shipshape.

> Practical guides on data foundations, governance and AI delivery from Shipshape Data. Written by the team that builds these systems for clients.

Canonical page: https://shipshapedata.com/resources/

On this page:

- Data & architecture
  - Data cleaning: what it involves
  - Data contracts: what they are and how they work
  - What is a data dictionary?
  - Data ingestion explained: batch, streaming, and CDC
  - Data orchestration: how data pipelines get scheduled, retried, and rerun
  - Data reconciliation: why your systems disagree and how to fix it
  - Data warehousing: how it works and when you need it
  - Star schema vs snowflake schema: choosing the right dimensional model
  - Batch versus real-time processing: choosing how data moves
  - 10 best data lineage tools for enterprise (2026)
  - 12 open-source data pipeline tools for modern ETL (2026)
  - Cloud migration explained: strategy, phases, steps, and costs
  - Data lake vs data warehouse: key differences for AI, ML, and BI
  - 7 data pipeline design patterns for modern analytics (2026)
  - ETL vs ELT: key differences, performance, and use cases
  - What is a data lakehouse? Architecture, benefits, and use cases
  - What is data integration? Definition, types, and use cases
  - What is a data pipeline? Definition, types, and key examples
  - Data pipeline architecture: a complete guide
  - What is data engineering? Skills, tools, use cases, and ROI
  - Data lakehouse vs data warehouse: which architecture wins?
  - Data engineering vs data science: roles, skills, and salary
  - Collibra data quality: features, use cases, and benefits
  - 9 best data pipeline monitoring tools for enterprise teams
  - 12 open-source data observability tools compared (2026)
  - Data quality management: dimensions, frameworks, and practices
  - Data silos: definition, problems, and how to eliminate them
  - Data strategy: definition, pillars, frameworks, and steps
  - What is data lineage? Definition, benefits, and examples
  - The complete guide to Informatica data quality (IDQ/CDQ)
  - Data lakehouse architecture: definition, layers, and diagrams
  - Data quality: what it is, why it matters, and how to improve it
  - How to improve data quality: 7 practical, proven strategies
  - Tableau data lineage: what it is and how to trace impact
  - Cloud migration roadmap: a step-by-step plan for enterprises
  - Databricks data quality: tools, rules, and monitoring at scale
  - Data engineering consulting services: what they include
  - Bigeye data observability: features, use cases, and reviews
  - AWS data engineering: skills, services, and learning path
  - Azure data engineering: skills, tools, and career roadmap
  - Data quality framework: components, steps, and practices
  - Great Expectations data quality: how to automate validation
  - Monte Carlo data observability: features, pricing, and reviews
  - Data quality dimensions: definitions, examples, and metrics
  - Soda data quality: how to implement, integrate, and compare
- AI & integration
  - RAG architecture: how to build a pipeline that holds up
  - AI due diligence: what private equity investors need to check
  - Predictive maintenance: how it works and why it matters
  - AI sales assistants: what they do and where they break
  - Dynamic content optimisation: how it works and where it goes wrong
  - Latency in AI systems: why responses take as long as they do
  - AI workflow orchestration: coordinating the steps behind an AI feature
  - Fine-tuning: when it's the right tool, and when it's the expensive one
  - Context windows: how much a language model can consider at once
  - Multimodal interactions: text, images, audio, and documents, working together
  - Agentic AI: what it is and when it is worth building
  - Preventing hallucination in AI systems
  - Conversational AI: chat interfaces over your own data and processes
  - Feature engineering: turning raw data into the inputs a model can learn from
  - Model drift: why a model that worked at launch degrades over time
  - Model validation and testing: how you know it's ready to deploy
  - What is model deployment? From training to production
  - Generative AI agents: definition, architecture, and use cases
  - MLOps architecture: components, design principles, and workflow
  - What is MLOps? Benefits, lifecycle, and best practices
  - What is prompt engineering? Techniques, examples, and tips
  - Retrieval-augmented generation: how RAG works for teams
  - What is a vector database? How embeddings power AI search
  - Embeddings: how they work, APIs, and enterprise use cases
  - Generative AI explained: what it is, with examples
  - Feature store in MLOps: what it is and why it matters
  - Qdrant vector database: architecture, indexing, and scale
  - Chroma vector database: features, setup, and use cases
  - Pinecone vector database: features, use cases, and setup
  - Chain of thought prompting: definition, examples, and templates
  - How to use OpenAI Evals to test and tune LLMs in production
  - How to evaluate large language models: metrics and benchmarks
  - Model monitoring in production: metrics, drift, and alerts
  - 13 prompt engineering best practices for better LLM output
  - 12 best MLOps tools for tracking, deployment, and monitoring
  - Meta Llama model access: official registration and platforms
  - Databricks LLM evaluation: MLflow metrics and best practices
  - AWS SageMaker model monitoring: how to set up drift alerts
  - Cohere prompt engineering: techniques, tips, and examples
  - Anthropic prompt engineering guide: master Claude models
  - Prompt engineering: a practical guide with tips and examples
- AI foundations
  - Industrial AI: what it is and where it works
  - Smart manufacturing: what it means and how to start
  - Artificial general intelligence: what AGI would mean and what it doesn't change yet
  - Deep learning: what more layers buy you
  - Natural language processing: getting computers to work with human language
  - Machine learning: how systems learn patterns from data
  - AI augmentation: extending what people can do rather than replacing them
  - Federated learning: how it works and when it's the right answer
  - Latent space: what it is and why it is hard to read
  - AI democratisation: widening access without losing control
  - AI as a service: what you're buying
  - Large language models: what an LLM is and how it works
  - The AI workforce: how AI changes the shape of work and teams
  - Artificial intelligence: a grounded definition for a business audience
  - Ensemble learning: bagging, boosting, and stacking explained
  - AI implementation strategy: from first use case to live
  - AI maturity models: what each stage looks like
  - AI-driven decision support: what it is and how to build it
  - Inference in AI: what happens when a model runs
  - Neural networks: how they work, with simple business examples
  - What is supervised learning? Basics, algorithms, and examples
  - Unsupervised learning: a complete guide for business
  - Reinforcement learning: definition, how it works, and examples
  - Overfitting: what it is, why it happens, and how to fix it
  - Hyperparameters: a plain-English guide to tuning
  - Predictive analytics: what it is and how it works, with examples
  - Enterprise AI: what it is, benefits, use cases, and platforms
  - Narrow AI: definition, examples, and enterprise use cases
- Unstructured data
  - Data augmentation: manufacturing useful variety when you cannot get more real data
  - Unstructured data: definition, examples, and AI use cases
  - What is a knowledge graph? Turning data into context
  - Synthetic data: definition, generation, use cases, and privacy
- AI governance & ethics
  - EU AI Act summary: risk tiers, obligations, and timeline
  - Black box AI: why models stay opaque and how to reduce it
  - Human in the loop: keeping people in AI decision processes
  - What is AI governance? Principles, risks, and how to start
  - What is responsible AI? Principles, practices, and governance
  - Ethical AI: principles, governance, and practical steps
  - AI bias: causes, types, examples, and how to reduce it
  - The complete guide to the NIST AI Risk Management Framework
  - Explainability in AI: what it is and why it builds trust
  - OECD AI principles: the five values for trustworthy AI
  - AI governance checklist: 12 steps for enterprise AI teams
  - AI governance vs data governance: key differences and overlap
  - Responsible AI governance framework: principles and steps
  - AI governance best practices: framework, policies, and controls
  - AI risk management: NIST AI RMF, steps, and best practices
  - AI governance framework: from principles to implementation
  - Model cards: what they are, examples, and how to create one
- Business intelligence
  - Demand forecasting: what it is and how to get it right
  - Inventory optimisation: holding the right stock, in the right place
  - Supply chain analytics: what it is and how to use it
  - Supply chain optimisation: what it means and how to approach it
  - A/B testing: comparing two versions properly
  - Business intelligence: what it is and how the stack fits together
  - MicroStrategy business intelligence: architecture explained
  - Business intelligence vs data analytics: a practical guide
  - Business intelligence architecture: basics, design, and examples
  - Self-service business intelligence: definition, tools, and tips
  - Qlik business intelligence: platform, features, and dashboards
  - Oracle business intelligence: features, editions, and pricing
  - 9 best business intelligence tools for enterprises in 2026
  - Attribution modelling: what it is, types, and how to choose
  - 11 best business intelligence software in 2026 (free and paid)
  - How to choose cloud business intelligence solutions
  - SAP business intelligence: what it is, tools, and comparisons
- Data governance & compliance
  - Compliance automation: controls as code and continuous evidence
  - Data privacy under GDPR and CCPA: a practical guide
  - Master data management: what it is and how to start
  - 16 master data governance solutions for enterprise (2026)
  - The DAMA data governance framework: principles and best practices
  - Managed data security services: what they are and what you get
  - 12 best data governance tools compared for enterprise (2026)
  - Data security governance framework: components and template
  - What is GDPR compliance? Principles, requirements, and checklist
  - Collibra data governance: features, benefits, and use cases
  - 12 data governance best practices to implement in 2026
  - Data governance consulting services: what to expect in 2026
  - What is a data protection impact assessment (DPIA)?
  - What is data security? Principles, types, and best practices
  - What is data governance? Pillars, roles, examples, and ROI
  - GDPR data subject rights: the eight rights and how to use them
  - Data quality governance: definition, models, and strategy

---

Shipshape Data is a London AI consultancy: we build the data foundation your AI depends on, then the AI on top. Site guide for agents: https://shipshapedata.com/llms.txt | Contact: hello@shipshapedata.com
