* [Topics](/en/topics "Topics")
* [Open source](/en/topics/open-source "Open source")
* What is InstructLab?
What is InstructLab?
====================
Published  October 6, 2025•*4*-minute read
Copy URL
Jump to section
---------------
What is InstructLab?How does InstructLab work?How is InstructLab different?What is Red Hat InstructLab on IBM Cloud?
What is InstructLab?
--------------------
InstructLab simplifies the process of customizing [large language models (LLMs)](/en/topics/ai/what-are-large-language-models) with private data.
LLMs serve as the foundation for [generative AI (gen AI)](/en/topics/ai/what-is-generative-ai), like chatbots and coding assistants. These LLMs can be proprietary (such as OpenAI’s GPT models and Anthropic’s Claude models) or offer varying degrees of openness around pretraining data and usage restrictions (such as Meta’s Llama models, Mistral AI’s Mistral models, and [IBM’s Granite models](http://redhat.com/en/topics/ai/what-are-granite-models)).
AI practitioners often need to adapt a pretrained LLM to suit a particular business purpose. But there are limits to the ways you can modify an LLM:
* Fine-tuning an LLM to understand a specific area of knowledge or skills typically involves expensive, resource-intensive training.
* There’s no way to incorporate improvements back to the LLM, and thus no way for models to continuously improve from user contributions.
* LLM refinements have typically required large amounts of human-generated data, which can be time-consuming and expensive to get.
InstructLab follows an approach that punches through those limitations. It can enhance an LLM using far less human-generated information and far fewer computing resources than are typically used to retrain a model.
InstructLab is named after and based on IBM Research’s work on Large-scale Alignment for chatBots, abbreviated as LAB. The LAB method is described in a [2024 research paper](https://arxiv.org/abs/2403.01081) by members of the MIT-IBM Watson AI Lab and IBM Research.
InstructLab is not model-specific. It can provide supplemental skills and knowledge fine-tuning to an LLM of your choice. This “tree of skills and knowledge” improves continuously from contributions and can be applied to support regular builds of an enhanced LLM.
The InstructLab project prioritizes fast iteration and intends to retrain models on a regular basis. Organizations can also use the InstructLab model alignment tools to train their own private LLMs with their own proprietary skills and knowledge.
[How to apply AI at the enterprise](/en/topics/ai/what-is-enterprise-ai)
How does InstructLab work?
--------------------------
The LAB method consists of 3 components:
* **Taxonomy-driven data curation.** Taxonomy is a set of diverse training data curated by humans as examples of new knowledge and skills for the model.
* **Large-scale synthetic data generation.** The model is then used to generate new examples based on the seed training data. Recognizing that synthetic data can vary in quality, the LAB method adds an automated step to refine the example answers, making sure they’re grounded and safe.
* **Iterative, large-scale alignment tuning.** Finally, the model is retrained based on the set of synthetic data. The LAB method includes 2 tuning phases: knowledge tuning, followed by skill tuning.
The dataset grows through contributions of skills and knowledge, with each addition improving the quality of the model being tuned.
Recommended for you
Unlocking Advocacy: Red Hat Programs for Building Expertise, Networking, and Thought Leadership
-----------------------------------------------------------------------------------------------
[Watch the webinar](https://www.redhat.com/en/events/webinar/unlocking-advocacy-redhat-programs-for-building-expertise-networking-and-thought-leadership?percmp=RHCTG0250000455235)
How is InstructLab different from other methods of training an LLM?
-------------------------------------------------------------------
Let’s compare InstructLab to the other ways of customizing an LLM to serve your domain-specific use cases.
### Pretraining
Pretraining represents the deepest level of model customization, and allows organizations to curate what training data shapes the models’ foundational understanding.
During pretraining, an LLM is trained to predict the next token using trillions of tokens of unlabeled data. This gets really expensive, sometimes requiring thousands of GPUs and months of time. Pretraining a highly capable LLM is only possible for organizations with significant resources.
### Alignment tuning
After pretraining, LLMs undergo alignment tuning to make the model’s answers as accurate and useful as possible. The 1st step in alignment tuning is typically instruction tuning, in which a model is trained directly on specific tasks of interest.
Next is preference tuning, which can include reinforcement learning from human feedback (RLHF). In this step, humans test the model and rate its output, noting if the model’s answers are preferred or unpreferred. An RLHF process may include multiple rounds of feedback and refinement to optimize a model.
Researchers have found that the amount of feedback at this alignment tuning stage can be much smaller than the initial set of training data―tens of thousands of human annotations, compared to the trillions of tokens of data required for pretraining―and still unlock latent capabilities of the model.
### InstructLab
The LAB method emerged from the idea that it should be possible to realize the benefits of model alignment from an even smaller set of human-generated data. An AI model can use a handful of human examples to generate a large amount of synthetic data―then refine that list for quality―and use that high-quality synthetic data set for further tuning and training. In contrast to instruction tuning, which typically needs thousands of examples of human feedback, LAB can make a model significantly better using relatively few examples provided by humans.
### How is InstructLab different from retrieval-augmented generation (RAG)?
The short answer is InstructLab and [retrieval-augmented generation (RAG)](/en/topics/ai/what-is-retrieval-augmented-generation) solve different problems.
RAG is a cost-efficient method for supplementing an LLM with domain-specific knowledge that wasn’t part of its pretraining. RAG makes it possible for a chatbot to accurately answer questions related to a specific field or business without retraining the model.
With RAG, knowledge documents are stored in a vector database, then retrieved in chunks and sent to the model as part of user queries. This is helpful for anyone who wants to add proprietary data to an LLM without giving up control of their information, or who needs an LLM to access timely information.   
This is in contrast to the InstructLab method, which sources end-user contributions to support regular builds of an enhanced version of an LLM. InstructLab helps add knowledge and unlock new skills of an LLM.
It’s possible to "supercharge" a RAG process by using the RAG technique on an InstructLab-tuned model.
[Learn more about RAG](/en/topics/ai/what-is-retrieval-augmented-generation)
What is Red Hat InstructLab on IBM Cloud?
-----------------------------------------
[Red Hat AI InstructLab on IBM Cloud](https://www.redhat.com/en/products/ai/instructlab-on-ibm-cloud) allows for contributions to an LLM without the need to own and operate hardware infrastructure.
On its own, Red Hat InstructLab is an open source project that simplifies LLM development, enabling a cost-effective approach to align models with less data and fewer resources.
IBM Cloud is an enterprise cloud platform designed for even the most regulated industries, delivering a highly resilient, performant, secure, and compliant cloud.
Together, Red Hat AI InstructLab on IBM Cloud offers a scalable, cost-effective solution to simplify, scale, and secure the alignment of your LLMs to your organization’s unique use-cases.
[Learn more about Red Hat InstructLab on IBM Cloud](https://www.redhat.com/en/products/ai/instructlab-on-ibm-cloud)
Recommended for you
E-book
Solve business and IT challenges with Red Hat Enterprise Linux 10
-----------------------------------------------------------------
New features and capabilities in Red Hat Enterprise Linux 10 can help you solve key business and IT challenges. Download to read more.
[Read the e-book](https://www.redhat.com/en/resources/solve-business-challenges-rhel-10-ebook?percmp=RHCTG0250000455234)
Recommended for you
Red Hat Enterprise Linux AI Technical Overview
----------------------------------------------
This free Technical Overview describes Red Hat Enterprise Linux AI (RHEL AI), a foundation model platform used to seamlessly develop, test, and run Granite family large language models (LLMs) for enterprise applications.
[Learn more](https://www.redhat.com/en/services/training/ai096-red-hat-enterprise-Linux-ai-technical-overview?percmp=RHCTG0250000455236)
Keep reading
------------
### What is Istio?
Find out more about Istio, an open source service mesh that controls how microservices share data with one another.
[Read the article](/en/topics/microservices/what-is-istio "article | What is Istio?")
### What is CentOS Stream?
CentOS Stream is a Linux® development platform where open source community members can contribute to Red Hat® Enterprise Linux in tandem with Red Hat developers.
[Read the article](/en/topics/linux/what-is-centos-stream "product article | what is centos stream")
### What is KVM?
Kernel-based virtual machines (KVM) are an open source virtualization technology that turns Linux into a hypervisor.
[Read the article](/en/topics/virtualization/what-is-KVM "article | what is KVM?")
Open source resources
---------------------
### Featured product
* #### [Red Hat AI](/en/products/ai)
[See all products](/en/technologies/all-products "See all products")
### Related content
* E-book
  [Unlock the power of agentic AI](/en/resources/executive-guide-ai-agentic-ebook)
* E-book
  [Before you build: A look at AI agentic systems with Red Hat AI](/en/resources/ai-agentic-systems-ebook)
* Blog post
  [The MLOps Challenge: Scaling from one model to thousands](/en/blog/mlops-challenge-scaling-one-model-thousands)
* Blog post
  [4 AI use cases for telco: Red Hat's autonomous intelligent network solutions](/en/blog/4-ai-use-cases-telco-red-hats-autonomous-intelligent-network-solutions)
### Related articles
* [What is distributed inference?](/en/topics/ai/what-is-distributed-inference)
* [What is Model Context Protocol (MCP)?](/en/topics/ai/what-is-model-context-protocol-mcp)
* [AIOps explained](/en/topics/ai/what-is-aiops)
* [What is AI security?](/en/topics/ai/what-is-ai-security)
* [Agentic AI vs. generative AI](/en/topics/ai/agentic-ai-vs-generative-ai)
* [What are large language models?](/en/topics/ai/what-are-large-language-models)
* [What is Models-as-a-Service?](/en/topics/ai/what-is-models-as-a-service)
* [What is AI in the public sector?](/en/topics/ai/what-is-ai-in-the-public-sector)
* [SLMs vs LLMs: What are small language models?](/en/topics/ai/llm-vs-slm)
* [What is enterprise AI?](/en/topics/ai/what-is-enterprise-ai)
* [What is Istio?](/en/topics/microservices/what-is-istio)
* [What is parameter-efficient fine-tuning (PEFT)?](/en/topics/ai/what-is-peft)
* [Why choose Red Hat Ansible Automation Platform as your AI foundation?](/en/topics/automation/automation-and-ai)
* [LoRA vs. QLoRA](/en/topics/ai/lora-vs-qlora)
* [What is CentOS Stream?](/en/topics/linux/what-is-centos-stream)
* [What is vLLM?](/en/topics/ai/what-is-vllm)
* [What is AI inference?](/en/topics/ai/what-is-ai-inference)
* [Predictive AI vs generative AI](/en/topics/ai/predictive-ai-vs-generative-ai)
* [What is agentic AI?](/en/topics/ai/what-is-agentic-ai)
* [What is KVM?](/en/topics/virtualization/what-is-KVM)
* [What are Granite models?](/en/topics/ai/what-are-granite-models)
* [RAG vs. fine-tuning](/en/topics/ai/rag-vs-fine-tuning)
* [What is Podman Desktop?](/en/topics/containers/what-is-podman-desktop)
* [Understanding AI in telecommunications with Red Hat](/en/topics/ai/understanding-ai-in-telecommunications)
* [Edge solutions for real-time decision making](/en/topics/edge-computing/edge-solutions-real-time-decision-making)
* [What are CentOS replacements?](/en/topics/linux/centos-alternatives)
* [What is CentOS?](/en/topics/linux/what-is-centos)
* [What are intelligent applications?](/en/topics/ai/what-are-intelligent-applications)
* [What is Podman?](/en/topics/containers/what-is-podman)
* [What is retrieval-augmented generation?](/en/topics/ai/what-is-retrieval-augmented-generation)
* [What is Helm?](/en/topics/devops/what-is-helm)
* [What is Argo CD?](/en/topics/devops/what-is-argocd)
* [What is an AI platform?](/en/topics/ai/what-is-an-ai-platform)
* [What is LLMops](/en/topics/ai/llmops)
* [What are predictive analytics](/en/topics/automation/how-predictive-analytics-improve-it-performance)
* [What is deep learning?](/en/topics/ai/what-is-deep-learning)
* [AI in banking](/en/topics/ai/ai-in-banking)
* [What is MicroShift?](/en/topics/edge-computing/microshift)
* [AI infrastructure explained](/en/topics/ai/ai-infrastructure-explained)
* [Understanding AI/ML use cases](/en/topics/ai/ai-ml-use-cases)
* [What is MLOps?](/en/topics/ai/what-is-mlops)
* [What are foundation models for AI?](/en/topics/ai/what-are-foundation-models)
* [How Kubernetes can help AI/ML](/en/topics/cloud-computing/how-kubernetes-can-help-ai)
* [What is generative AI?](/en/topics/ai/what-is-generative-ai)
* [What is edge AI?](/en/topics/edge-computing/what-is-edge-ai)
* [OpenJDK versus Oracle JDK](/en/topics/application-modernization/openjdk-vs-oracle-jdk)
* [What is Cloud Foundry?](/en/topics/application-modernization/what-is-cloud-foundry)
* [What is Kubeflow?](/en/topics/cloud-computing/what-is-kubeflow)
* [What is Buildah?](/en/topics/containers/what-is-buildah)
* [Accelerate MLOps with Red Hat OpenShift](/en/technologies/cloud-computing/openshift/aiml)
* [Understanding Ansible, Terraform, Puppet, Chef, and Salt](/en/topics/automation/understanding-ansible-vs-terraform-puppet-chef-and-salt)
* [Ansible vs. Chef: What you need to know](/en/topics/automation/ansible-vs-chef)
* [Ansible vs. Salt: What you need to know](/en/topics/automation/ansible-vs-salt)
* [What is Linux?](/en/topics/linux/what-is-linux)
* [What's the best Linux distro for you?](/en/topics/linux/whats-the-best-linux-distro-for-you)
* [What is machine learning?](/en/topics/ai/what-is-machine-learning)
* [Ansible vs. Puppet: What you need to know](/en/topics/automation/ansible-vs-puppet)
* [Red Hat OpenShift vs. OKD](/en/topics/containers/red-hat-openshift-okd)
* [What is AI in healthcare?](/en/topics/ai/what-is-ai-in-healthcare)
* [What is Apache Kafka?](/en/topics/integration/what-is-apache-kafka)
* [Spring on Kubernetes with Red Hat OpenShift](/en/technologies/cloud-computing/openshift/spring)
* [Why run Apache Kafka on Kubernetes?](/en/topics/integration/why-run-apache-kafka-on-kubernetes)
* [Ansible vs. Terraform, clarified](/en/topics/automation/ansible-vs-terraform)
* [Ansible vs. Red Hat Ansible Automation Platform](/en/technologies/management/ansible/ansible-vs-red-hat-ansible-automation-platform)
* [What is Skopeo?](/en/topics/containers/what-is-skopeo)
* [Using Helm with Red Hat OpenShift](/en/technologies/cloud-computing/openshift/helm)
* [What is Grafana?](/en/topics/data-services/what-is-grafana)
* [What is open source software?](/en/topics/open-source/what-is-open-source-software)
* [Open source vs. proprietary software in vehicles](/en/topics/open-source/open-source-vs-proprietary-software-in-vehicles)
* [What is KubeLinter?](/en/topics/containers/what-is-kubelinter)
* [What is RKT?](/en/topics/containers/what-is-rkt)
* [What is Kogito?](/en/topics/automation/what-is-kogito)
* [What was CoreOS and CoreOS container Linux](/en/technologies/cloud-computing/openshift/what-was-coreos)
* [What is Jaeger?](/en/topics/microservices/what-is-jaeger)
* [What is open source?](/en/topics/open-source/what-is-open-source)
* [What is Clair?](/en/topics/containers/what-is-clair)
* [What is Knative?](/en/topics/microservices/what-is-knative)
* [What is etcd?](/en/topics/containers/what-is-etcd)
* [What is Docker?](/en/topics/containers/what-is-docker)
[More about this topic](/en/topics/open-source "More about this topic")