Celeste AICeleste AI
Core concepts

Modality & operations

The core mental model of Celeste v1.

Architecture & Concepts

Celeste v1 is built around a simple mental model: Domain (what you work with) + Operation → Modality (output).

Domain
What you work with

Images, Audio, Video, Text

Namespace API

Modality
What comes out

Output type and client

create_client(modality=...)

Operation
How input becomes output
GeneratePrompt → domain outputoutput = modality
EditDomain → same domainimages client
AnalyzeMedia → texttext client
EmbedAny domain → vectorembeddings client
Namespaces use domainsceleste.images.analyze(...)
Clients use modalitiescreate_client(modality="text")
Image generation → domain: images, output modality: images
Image analysis → domain: images, output modality: text

The Core Trinity

1. Modality (The output)

What comes out. Modality defines the output type you receive: text, images, audio, video, or embeddings.

2. Domain (What you work with)

The resource you operate on. Domain is not input vs output; it is the thing you are working with. That resource can be input or output depending on the operation.

  • Image generation: domain = images, output modality = images.
  • Image analysis: domain = images, output modality = text.

3. Operation (The transformation)

The method you call. Operations describe how the domain is transformed into the output modality.

Operation Matrix (Namespace-first)

This matrix explains where to find every operation in Celeste.

What you work with (Domain)OperationOutput (Modality)Namespace call
TextgenerateTextawait celeste.text.generate(...)
TextembedEmbeddingsawait celeste.text.embed(...)
ImagesgenerateImagesawait celeste.images.generate(...)
ImageseditImagesawait celeste.images.edit(...)
ImagesanalyzeTextawait celeste.images.analyze(...)
AudiospeakAudioawait celeste.audio.speak(...)
AudioanalyzeTextawait celeste.audio.analyze(...)
VideosgenerateVideosawait celeste.videos.generate(...)
VideosanalyzeTextawait celeste.videos.analyze(...)

Operations by Domain

Actiontextimagesvideosaudio
generate✓✓✓○
edit·✓··
analyze·✓✓✓
upscale·○○·
speak···✓
transcribe···○
embed✓○○·

Legend: ✓ = Available · ○ = Planned · · = Not applicable

Note: This table is by domain (what you work with). Embeddings is a modality (vector output) produced by the embed operation, so it is not listed as a domain.

Examples

Image Generation (Image domain → Image output)

import celeste

response = await celeste.images.generate(
    model="gpt-image-1",
    prompt="A red apple",
)
image = response.content

Image Analysis (Image domain → Text output)

import celeste
from celeste.artifacts import ImageArtifact

image = ImageArtifact(path="apple.png")

response = await celeste.images.analyze(
    model="gpt-4o",
    prompt="Describe this image",
    image=image,
)
text = response.content

Why this design?

Celeste keeps outputs predictable: a text response always returns TextOutput, an image response always returns ImageOutput, and so on. The namespace-first API mirrors how people think (“I have an image; I want to analyze it”) while the modality system keeps output types consistent and type-safe.