We can LLMs in three ways by usually
AI agents are semi autonomous systems that interact with environment, make decisions and perform tasks on behalf of users.
Cursor 2.0 is an AI editor for Production Environment. It will be run 8 parallel agents without any issue.
Context Management like telling story when it getting convoluted. It will direct path when AI get confused.
Context window is a windows chat where user and AI interact each other.
LLM stands for Large Language Model. It is specifically a deep learning model, trained on massive amounts of text data to understand and generate human language, enabling tasks like text generation, translation. It often sing "Transformer" models which are neural networks that can process relationships within language.
AI - System or machines that mimic human intelligence to perform tasks and can iteratively improve themselves based on the information they collect. Artificial Intelligence capable of generating text, images, videos or other data using generative models often in response to prompts. Generative AI models learn the patterns and structure of their input training data and then generate new data that has similar characteristics.
AGI - Artificial General Intelligence is a type of AI that can understand, learn and apply knowledge across broad range of tasks similar to human cognitive abilities.
Example of Cloud Machine Learning:
Different AI Types:
Machine Learning:
Simple Input -> Simple output -> Single topic
Deep Learning:
Complex input -> Simple output -> Single topic
Foundation Model:
Complex inputs -> Complex ouputs -> multiple topics
Foundation Model:
A foundation model is a type of large scale artificial intelligence
model that is trained on a broad range of data at massive scale, allowing to
develop a wide understanding of many topics and tasks. These models can be adapted or fine-tuned for
various specific applications, demonstrating flexibility and efficiency across
different domains.
Parameters of foundation models:
LLM:
Encoders:
Models that convert a sequence of words to an embedding.
Decoders:
Models take a sequence of words and output next word.
Examples: GPT-4, Llama and bloom
Encoder - decoder module:
We passed english letters and encoder covert into token. Decode passed one tocken at time.
Hallucination:
It is generated text that is non factual and Or ungrounded.
LLM application:
Retrieval Augmented Generation (RAG)
Code models:
In-context learning and few shot prompting:
Language Agents:
* A Budding area of research where LLM based agents
Some notable work in the space:
* ReAct
Iterative framework where LLM emits thoughts, then act and observes result
* Toolformer
Pre-training technique where strings are replaced with calls to tools that yield result.
OCI Generative AI service:
* Fully managed service that provides a set of customizable Large Language Models (LLM) avilable via a single API to build generative AI applications.
Generation:
Command -> Command light -> llama 2.7
Dedicated AI cluster:
* Dedciated AI cluster has a GPU based resource that host the customers fine-tuning and inference workloads.
OCI setup:
configuration file : ./oci/config
Model parameters:
Temperature : Determines how creative module should be, default temperature is 1 and maximum temperature is 5.
Length: Approximate length of the summary, choose from short, medium and length.
Embeddings:
Embedding is a numerical represent of piece of text converted into number sequences.
A piece of text could be a word, phrase, sentence or paragraph or more paragraphs.
Models creates a 1024 vector for each embedding.
Max 512 tokens per embedding.
Model create a 384 dimensional vector for each embedding.
It will be very difficult to tune a 2 billion tokens. We are using the In-context Learning/Few shot Prompting:
GPU memory is limited, so switching between models can incur significant overhead due to reloading the full GPU memory.
Dedicated AI cluster units:
* Large cohere - Dedicated AI cluster units for hosting or fine tuning the cohere command
* Small Cohere - Dedicated AI cluster units for hosting or fine tuning for small cohere command
* Embed Cohere - Dediated AI cluster for hosting the models
* Llama2-70 model - Dedicated AI cluster for hosting the Llamba2 models
Fine tunnning is required 2 units and each cluster is active for five hours.
RAG framework:
Retriever : It is act like search engine.
Ranker : Evalate and priorites rank based a quality of the data.
Generator : It provide human like texts.
RAG techniques:
RAG sequence
RAG token
RAG pipelines:
Documents -> Chunks -> Embedding -> Index [database]
Vector database:
A vector is sequence of numbers called dimensions, used to capture the important "features" of the data.
Semantic search:
It means search by meaning rather than giving a number.