Running AI Models with Hugging Face: A Deep Dive
What is Hugging Face?
Hugging Face is a platform built to help people use advanced AI models. It provides access to thousands of pre-trained models for various AI tasks like language processing, text generation, computer vision, audio analysis, and more.
You can think of Hugging Face as a library and marketplace for AI models, where people easily share, discover, and use powerful AI technology.
How Hugging Face works?
Hugging Face provides:
Model Hub: A central place where thousands of pre-trained models are hosted.
Transformers Library: A Python library allowing you to load and use models easily in your code.
Datasets & Spaces: Tools to build, share, and deploy AI apps quickly.
Running a Hugging Face model (Step-by-Step):
Here is how you typically run a Hugging Face model:
Step 1: Create a Hugging Face account and get your Token
Sign up at huggingface.co.
Go to Settings → Access Tokens.
Create a new token with necessary permissions (read-only usually enough).
Keep this token private—treat it like a password.
Important:
Always respect model licenses and usage rules. Many models have specific terms of use, so get consent or approval if needed before using any Hugging Face model publicly or commercially.
Step 2: Google Colab Notebook to run a Hugging Face model
Google Colab is a free environment provided by Google where you can easily run Python code and AI models. It’s an excellent platform to start exploring Hugging Face.
Below is a cleaned and improved version of your provided Colab notebook code.
✅ Cleaned Google Colab Notebook example
# Step 1: Install transformers and torch
!pip install transformers torch
# Step 2: Import libraries
import os
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
# Step 3: Set your Hugging Face token securely (replace 'your-token-here')
os.environ["HF_TOKEN"] = "your-token-here"
# Step 4: Define the model name
model_name = "google/gemma-3-1b-it"
# Step 5: Load tokenizer and model
tokenizer = AutoTokenizer.from_pretrained(model_name, token=os.environ["HF_TOKEN"])
model = AutoModelForCausalLM.from_pretrained(
model_name,
torch_dtype=torch.bfloat16,
token=os.environ["HF_TOKEN"]
)
# Step 6: Create your input prompt
input_prompt = "The capital of India is"
# Step 7: Tokenize the input prompt
tokenized_input = tokenizer(input_prompt, return_tensors="pt")
# Step 8: Generate output from the model
generated_tokens = model.generate(**tokenized_input, max_new_tokens=25)
# Step 9: Decode the output tokens to text
output_text = tokenizer.decode(generated_tokens[0], skip_special_tokens=True)
print(output_text)
How this works in simple terms:
Tokenizer breaks your input prompt into tokens (words or parts of words) the model can understand.
Model takes these tokens and generates new tokens to complete the prompt.
The generated tokens are then converted back into readable text.
Running this Code on Your Local Machine
You can run the above code directly on your own computer. However, we use Google Colab here because most powerful AI models require GPUs, which are specialized and expensive hardware. Google Colab provides free GPU access, while typical personal computers often do not have this capability.
Hugging Face vs Ollama: What's the difference?
1. Purpose & Complexity:
Ollama:
Simple, easy setup specifically designed to run AI models locally.
Provides an easy-to-use API to interact quickly with models.
Minimal configuration needed, ideal for beginners.
Hugging Face:
Broader and more powerful, giving detailed access and flexibility.
Requires more technical understanding to configure and run models.
Good for customized and research-oriented use cases.
2. Deployment Approach:
Ollama:
Simplified API access, managed deployment.
Runs models in container-like environments (similar to Docker).
Hugging Face:
Directly integrates into Python apps via libraries.
Can be deployed anywhere (local machine, cloud, web apps).
3. Model Availability:
Ollama:
- Limited, curated models optimized for ease of use and local deployment.
Hugging Face:
- Huge variety (thousands of models), covering many AI tasks and domains.
About Google Colab:
Google Colab is a free, cloud-hosted Python environment by Google, great for beginners and researchers.
Benefits:
No setup required; instantly usable.
Free access to GPUs (limited daily usage).
Integrates directly with Hugging Face easily.
Important Note About Consent:
When using any Hugging Face model, always:
Check licenses: Make sure you have permission to use the model.
Generate Token Carefully: Keep your Hugging Face tokens secure and private.
Respect user privacy & compliance: Ensure you have proper consent to use datasets and models in your projects, especially if they involve public or sensitive data.
Summary (Simplified):
Hugging Face: Large, flexible, community-driven, ideal for detailed and complex AI needs.
Ollama: Simple, beginner-friendly local model deployment, easy to use with minimal technical effort.
Google Colab: A great platform to quickly experiment with Hugging Face without needing any local setup.
