[ Blog ]

A map of machine learning: the task categories companies need to know

  • 11 September 2026
  • 5 min read

When managers think about AI, most think of chat apps: ChatGPT, Claude and company. But chat apps are one corner of a much larger map. Machine learning is divided into dozens of task categories, and each of them solves its own kind of problem. A company that knows the map does not ask "what can ChatGPT do for us" but "which category solves this task". That question opens far more doors.

The map is organised by input and output

The simplest way to understand the map is to ask two questions about every task: what goes in, and what should come out? Text in and text out is one task. An image in and a list of objects out is another. Speech in and written text out is a third. Each combination of input and output is a separate task category with its own models, its own datasets and its own best solutions.

The category names are deliberately kept in their original form here. They are the search terms on Hugging Face, the world's largest open model library, where each category lists both models and datasets. Anyone who knows the names can look up for themselves what exists, what is open and what is usable.

Natural Language Processing: everything to do with text

The text categories are the ones most people know best, but there is more to them than chat:

  • Text Generation: the model writes text from instructions. This is the category chat apps are built on.
  • Summarization: a long document in, a short summary out. Meeting minutes, reports, articles.
  • Translation: text in one language in, the same content in another language out.
  • Text Classification and Token Classification: text gets a label. A whole sentence is classified as a complaint or praise, or individual words are tagged as names, places and amounts.
  • Question Answering: a question and a document in, an answer out. There are versions that answer questions straight from tables.
  • Zero-Shot Classification: sorting into categories the model has never seen before, without special training.
  • Sentence Similarity: how close are two sentences in meaning? The foundation of semantic search.
  • Text Ranking: orders documents by how well they answer a query.
  • Feature Extraction: turns text into numerical vectors that other systems can work with.
  • Fill-Mask: the model fills in gaps in text, the basic task behind many language-understanding models.

Computer Vision: everything to do with images

The image categories are more varied than many people realise:

  • Image Classification: what is in the picture? One label per image.
  • Object Detection: which objects are in the picture, and where? A box around each object.
  • Image Segmentation: which pixels belong to which object? A more precise version of the same question.
  • Text-to-Image and Image-to-Image: an image is created from text, or one image is transformed into another.
  • Text-to-Video and Image-to-Video: the same idea, but the output is video.
  • Depth Estimation: how far is each part of the image from the camera?
  • Keypoint Detection: finds key points, for example people's joints or the corners of objects.
  • Mask Generation: creates precise outlines of objects, often with a single click.
  • Zero-Shot versions: classification and object detection without special training on the objects in question.
  • Text-to-3D and Image-to-3D: a 3D model is created from text or an image.

Multimodal: when the input is mixed

The most interesting categories for office work are often the ones that take more than one type of data at once:

  • Image-Text-to-Text: an image and a question in, a text answer out.
  • Visual Question Answering: a specialised version, where questions are asked directly about the content of an image.
  • Document Question Answering: questions are answered from documents with layout, tables and images, for example invoices or manuals.
  • Audio-Text-to-Text and Video-Text-to-Text: audio or video together with text in, text out.
  • Visual Document Retrieval: searching a document archive where the appearance of the documents matters, not just the text.
  • Any-to-Any: models that accept and return any combination.

Audio: everything to do with sound

  • Automatic Speech Recognition: speech in, written text out.
  • Text-to-Speech: text in, speech out.
  • Text-to-Audio: text in, sound or music out.
  • Audio Classification: what sound is this? Machine noise, an alarm, music.
  • Voice Activity Detection: at which points in a recording is someone speaking?

Tabular and time series: everything to do with tabular data

These categories get the least media attention but often deliver the most in operations:

  • Tabular Classification: predicts a category from tabular data, for example whether a customer is likely to churn.
  • Tabular Regression: predicts a number from tabular data, for example a price or a risk.
  • Time Series Forecasting: predicts future values from historical trends, for example sales in the coming weeks.

In addition, there are three categories most companies only need to know exist: Reinforcement Learning, where a model learns by trial and reward, Robotics, where models control hardware, and Graph Machine Learning, where the data is a network of relationships, for example business connections or supply chains.

What does the map mean for Icelandic companies?

The map only becomes useful once the categories are connected to real tasks. A few examples that apply to Icelandic companies:

  • Object Detection on shelf photos. A sales rep takes a photo of a shelf in a shop and the model counts which products are visible. Shelf share, previously estimated by counting on site, becomes measurable from a photo library.
  • Document Question Answering on internal manuals. Staff ask in plain language and get an answer straight from the quality manual, the safety rules or the collective agreement, with a reference to the right page.
  • Automatic Speech Recognition on meetings and podcasts. Recordings become searchable text. Last year's meeting, where the decision was made, is found in seconds.
  • Time Series Forecasting on seasonal demand. Tourism, retail and restaurants in Iceland swing with the seasons. A forecasting model on sales history improves purchasing, staffing and inventory.
  • Sentence Similarity on your own content. A search on your website or intranet that finds the right answer even when the wording differs from the document. "Can I get a refund" finds the document called "Returns policy".

Notice that none of these tasks is chat. All of them are still machine learning, and open models exist for all of them on Hugging Face. The right question is never "what can ChatGPT do" but "which category solves this task".

What you can do right now

  • Write down three tasks and analyse their input and output. Pick three time-consuming tasks in your business and note for each of them: what goes in, and what should come out? Compare the answers with the categories above.
  • Browse Hugging Face for 20 minutes. Go to huggingface.co, pick one category that matches your task and look at which models and datasets are there. You do not need to understand the technology to see what exists.
  • Ask where your data is sitting unused. Photos, recordings, sales history, document archives. Each type of data has its own categories on the map. Data nobody reads today can become a measuring instrument tomorrow.

[ Get in touch ]

Book a free assessment

90 minutes that pay off immediately: we map your AI usage, risks and 3 to 5 automatable workflows, and deliver a report within a week. No commitment.

No commitment

[ Direct contact ]

hallo@vestra.is+354 863 7496

Bolholt 8
105 Reykjavík, Iceland