#LLM
6 articles tagged LLM. Past halfway to the twelve-article threshold.
Timeline
-
What prompt caching is — from the second call, your input costs one-fortieth
-
What an AI token is — how much text $10 per million tokens actually buys
-
What an AI agent is — what changes when you hand a model tools
-
What fine-tuning is — and what prompting and RAG handle instead
All articles
-
Tech · 4 min readWhat prompt caching is — from the second call, your input costs one-fortieth
Mechanism — an identical prefix yields identical intermediate values, which can be stored
-
Tech · 4 min readWhat an AI token is — how much text $10 per million tokens actually buys
Unit — a token is a word fragment, not a word, and counts differ sharply by language
-
Tech · 4 min readWhat an AI agent is — what changes when you hand a model tools
Definition — software that takes a goal, picks tools, reads results and decides its own next step
-
Tech · 4 min readWhat fine-tuning is — and what prompting and RAG handle instead
Definition — continued training of an already-trained model on task data, changing its weights
-
Tech · 5 min readWhat AI token pricing is — why output costs about 5× input
Billing is per token, with separate input and output rates — output typically costs four to five times input
-
Tech · 4 min readWhat is a context window — a model does not read all million tokens
Context window = total tokens visible in one request, input and output combined