A.X 4.0 VL Light

๐Ÿค— Models | ๐Ÿ–ฅ๏ธ Github

Highlights

A.X 4.0 VL Light (pronounced โ€œA dot Xโ€) is a vision-language model (VLM) optimized for Korean vision and language understanding as well as enterprise deployment. Built upon A.X 4.0 Light, A.X 4.0 VL Light has been further trained on diverse multimodal datasets, with a particular focus on large-scale multimodal Korean datasets, to deliver exceptional performance in domestic business applications.

  • Superior Korean Proficiency in Vision and Language: Achieved an average score of 79.4 on Korean image benchmarks, outperforming Qwen2.5-VL-32B (73.4), despite having a significantly smaller model size. On Korean text benchmarks, recorded an average score of 60.2, comparable to VARCO-VISION-2.0-14B (60.4), while using only half the model size.
  • Deep Cultural Understanding: Scored 80.2 on K-Viscuit, a multimodal benchmark designed to evaluate cultural and contextual comprehension in Korean, exceeding Qwen2.5-VL-32B (72.3).
  • Advanced Document Understanding: Attained a score of 89.8 on KoBizDoc, a benchmark focused on understanding complex document structures, including charts and tables, performing comparably to Qwen2.5-VL-32B (88.8).
  • Efficient Token Usage: A.X 4.0 VL Light utilizes approximately 41% fewer text tokens compared to Qwen2.5-VL for the same Korean input, enabling significantly more cost-effective and efficient processing.

A brief comparison on representative benchmarks is as follows:

Performance

Image Benchmark

*Korean benchmarks, with K-Viscuit translated into Korean.

Category Benchmarks A.X 4.0 VL Light Qwen2.5-VL-7B InternVL3-8B VARCO-VISION-2.0-14B Qwen2.5-VL-32B
Document KoBizDoc* 89.8 84.0 73.2 83.0 88.8
K-DTCBench* 90.0 86.7 83.8 80.8 91.7
ChartQA 79.8 80.6 79.8 78.8 81.8
DocVQA 94.4 95.3 92.4 91.9 94.5
InfoVQA 78.5 82.7 76.2 80.0 82.7
SEEDBench2-Plus 69.7 71.2 69.7 71.9 73.3
OCR OutdoorKorean* 97.3 91.9 72.7 79.7 86.9
K-Handwriting* 84.3 85.0 43.5 55.2 60.1
TextVQA 82.0 85.4 82.1 80.3 79.8
Culture K-Viscuit* 80.2 65.0 65.3 72.0 72.3
Knowledge KoEduBench* 58.1 53.9 53.9 39.4 52.4
KoCertBench* 54.9 50.1 39.4 51.4 47.5
MMMU 54.1 56.3 59.4 58.3 63.6
ScienceQA 95.3 87.2 97.8 92.2 92.4
General K-LLaVA-W* 83.2 73.0 67.0 80.0 84.3
K-SEED* 76.5 76.4 76.4 76.9 77.3
SEEDBench_IMG 76.7 77.1 77.1 78.1 77.6
Hallucination HallusionBench 54.2 52.7 49.6 53.8 58.0
IF MM-IFEval 53.5 51.4 51.9 50.8 59.3

The following in-house benchmarks have been established to rigorously assess model performance on Korean vision-language understanding and the comprehension of Korea-specific knowledge domains:

  • KoBizDoc: A visual question answering (VQA) benchmark designed for understanding Korean business documents.
  • OutdoorKorean: A benchmark focused on recognizing Korean text in complex outdoor scenes (provided by AIHub).
  • K-Handwriting: A Korean handwriting recognition dataset comprising various handwritten styles (provided by AIHub).
  • KoEduBench: A VQA benchmark targeting Korean general academic exams, including GED and CSAT questions, to assess academic reasoning ability.
  • KoCertBench: A Korean certification exam-based VQA benchmark, covering domains such as civil service, technical licenses, and professional qualifications.

Text Benchmark

*Korean benchmarks.

Category Benchmarks A.X 4.0 VL Light Qwen2.5-VL-7B InternVL3-8B VARCO-VISION-2.0-14B
Knowledge KMMLU* 60.5 45.6 50.9 58.8
MMLU 72.6 71.9 77.5 80.7
Math HRM8K* 40.6 25.4 34.6 49.5
MATH 56.5 61.7 65.1 71.1
General Ko-MT-bench* 68.9 51.5 59.5 75.9
MT-bench 72.9 73.2 69.9 76.6
IF Ko-IFEval* 71.8 55.0 46.1 57.2
IFEval 81.9 66.6 67.5 75.3

๐Ÿš€ Quickstart

with HuggingFace Transformers

  • transformers>=4.49.0 or the latest version is required to use skt/A.X-4.0-VL-Light
pip install transformers>=4.49.0

Example Usage

import torch
from transformers import AutoModelForCausalLM, AutoProcessor
from PIL import Image
import requests
from io import BytesIO



model_name = "skt/A.X-4.0-VL-Light"
model = AutoModelForCausalLM.from_pretrained(model_name, trust_remote_code=True, torch_dtype=torch.bfloat16).to(device='cuda')
processor = AutoProcessor.from_pretrained(model_name, trust_remote_code=True)

url = "https://huggingface.co/skt/A.X-4.0-VL-Light/resolve/main/assets/image.png"
# ์ด๋ฏธ์ง€ ์ถœ์ฒ˜: ๊ตญ๊ฐ€์œ ์‚ฐํฌํ„ธ (https://www.heritage.go.kr/unisearch/images/national_treasure/thumb/2021042017434700.JPG)

response = requests.get(url)
response.raise_for_status()
image = Image.open(BytesIO(response.content))

messages = [
    {
        "role": "user",
        "content": [
            {"type": "image"},
            {"type": "text", "text": "์ด๋ฏธ์ง€์— ๋Œ€ํ•ด์„œ ์„ค๋ช…ํ•ด์ค˜."},
        ],
    }
]

text = processor.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True
)
inputs = processor(
    images=[image],
    text=[text],
    padding=True,
    return_tensors="pt",
).to("cuda")

# Decoding parameters (top_p, temperature, top_k, repetition_penalty) should be tuned depending on the generation task.
generation_kwargs = {
    "max_new_tokens": 256,
    "top_p": 0.8,
    "temperature": 0.5,
    "top_k": 20,
    "repetition_penalty": 1.05,
    "do_sample": True,
}
generated_ids = model.generate(**inputs, **generation_kwargs)
generated_ids_trimmed = [
    out_ids[len(in_ids) :] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
response = processor.batch_decode(
    generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
)
print(response[0])
"""
์ˆญ๋ก€๋ฌธ์€ ๋Œ€ํ•œ๋ฏผ๊ตญ ์„œ์šธ์— ์œ„์น˜ํ•œ ๊ตญ๋ณด ์ œ1ํ˜ธ๋กœ, ์กฐ์„  ์‹œ๋Œ€์— ๊ฑด์ถ•๋œ ๋ชฉ์กฐ ๊ฑด์ถ•๋ฌผ์ด๋‹ค. ์ด ๋ฌธ์€ ์„œ์šธ์˜ ๋‚จ์ชฝ ๋Œ€๋ฌธ์œผ๋กœ, ์ „ํ†ต์ ์ธ ํ•œ๊ตญ ๊ฑด์ถ• ์–‘์‹์„ ๋ณด์—ฌ์ค€๋‹ค. ๋‘ ์ธต์œผ๋กœ ์ด๋ฃจ์–ด์ง„ ์ด ๋ฌธ์€ ๊ธฐ์™€์ง€๋ถ•์„ ์–น๊ณ  ์žˆ์œผ๋ฉฐ, ์ง€๋ถ•์˜ ๊ณก์„ ์ด ์•„๋ฆ„๋‹ต๊ฒŒ ํ‘œํ˜„๋˜์–ด ์žˆ๋‹ค. ๋ฌธ ์•„๋ž˜์—๋Š” ์•„์น˜ํ˜•์˜ ์ถœ์ž…๊ตฌ๊ฐ€ ์žˆ์œผ๋ฉฐ, ๊ทธ ์ฃผ์œ„๋กœ๋Š” ๊ฒฌ๊ณ ํ•œ ์„์žฌ๋กœ ์Œ“์€ ์„ฑ๋ฒฝ์ด ์ด์–ด์ ธ ์žˆ๋‹ค. ๋ฐฐ๊ฒฝ์—๋Š” ํ˜„๋Œ€์ ์ธ ๊ณ ์ธต ๋นŒ๋”ฉ๋“ค์ด ์ž๋ฆฌ์žก๊ณ  ์žˆ์–ด, ์ „ํ†ต๊ณผ ํ˜„๋Œ€๊ฐ€ ๊ณต์กดํ•˜๋Š” ์„œ์šธ์˜ ๋ชจ์Šต์„ ์ž˜ ๋‚˜ํƒ€๋‚ธ๋‹ค. ์ˆญ๋ก€๋ฌธ์€ ์—ญ์‚ฌ์ , ๋ฌธํ™”์  ๊ฐ€์น˜๊ฐ€ ๋†’์•„ ๋งŽ์€ ๊ด€๊ด‘๊ฐ๋“ค์ด ์ฐพ๋Š” ๋ช…์†Œ์ด๋‹ค.
"""

Example for Document Transcription

import torch
from transformers import AutoModelForCausalLM, AutoProcessor
from PIL import Image
import requests
from io import BytesIO



model_name = "skt/A.X-4.0-VL-Light"
model = AutoModelForCausalLM.from_pretrained(model_name, trust_remote_code=True, torch_dtype=torch.bfloat16).to(device='cuda')
processor = AutoProcessor.from_pretrained(model_name, trust_remote_code=True)

url = "https://huggingface.co/skt/A.X-4.0-VL-Light/resolve/main/assets/document.png"

response = requests.get(url)
response.raise_for_status()
image = Image.open(BytesIO(response.content))

messages = [
    {
        "role": "user",
        "content": [
            {"type": "image"},
            {"type": "text", "text": "์‚ฌ์ง„์— ๋ฌด์—‡์ด ์ ํ˜€์žˆ๋‚˜์š”? ๋‹ค๋ฅธ ์„ค๋ช… ์—†์ด ์ ํ˜€์žˆ๋Š” ํ…์ŠคํŠธ๋งŒ ๊ฒฐ๊ณผ๋กœ ๋ณด์—ฌ์ค˜."},
        ],
    }
]

text = processor.apply_chat_template(
    messages, tokenize=False, add_generation_prompt=True
)
inputs = processor(
    images=[image],
    text=[text],
    padding=True,
    return_tensors="pt",
).to("cuda")


generation_kwargs = {
    "max_new_tokens": 1024,
    "top_p": 0.95,
    "top_k": 1,
    "temperature": 0.7,
    "repetition_penalty": 1.05,
    "do_sample": True,
}
generated_ids = model.generate(**inputs, **generation_kwargs)
generated_ids_trimmed = [
    out_ids[len(in_ids) :] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
response = processor.batch_decode(
    generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
)
print(response[0])
"""
# A.X 4.0: ๊ธฐ์—…์šฉ ํ•œ๊ตญ์–ด ํŠนํ™” ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ

View English README

SKํ…”๋ ˆ์ฝค์ด ํ•œ๊ตญ์–ด ์ฒ˜๋ฆฌ ๋Šฅ๋ ฅ๊ณผ ๊ธฐ์—… ํ™œ์šฉ์„ฑ์„ ๋†’์ธ ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM) A.X 4.0 (์—์ด๋‹ท์—‘์Šค 4.0)์„ 2025๋…„ 4์›” 30์ผ์— ์ถœ์‹œํ•˜์˜€์Šต๋‹ˆ๋‹ค. A.X 4.0์€ ์˜คํ”ˆ์†Œ์Šค ๋ชจ๋ธ์ธ Qwen2.5์— ๋ฐฉ๋Œ€ํ•œ ํ•œ๊ตญ์–ด ๋ฐ์ดํ„ฐ๋ฅผ ์ถ”๊ฐ€๋กœ ํ•™์Šต์‹œ์ผœ ๊ตญ๋‚ด ๋น„์ฆˆ๋‹ˆ์Šค ํ™˜๊ฒฝ์— ์ตœ์ ํ™”๋œ ์„ฑ๋Šฅ์„ ๋ฐœํœ˜ํ•ฉ๋‹ˆ๋‹ค.

## A.X 4.0, ๋ฌด์—‡์ด ๋‹ค๋ฅธ๊ฐ€์š”?

- ๋›ฐ์–ด๋‚œ ํ•œ๊ตญ์–ด ์‹ค๋ ฅ: ๋Œ€ํ‘œ์ ์ธ ํ•œ๊ตญ์–ด ๋Šฅ๋ ฅ ํ‰๊ฐ€ ๋ฒค์น˜๋งˆํฌ์ธ KMMLU์—์„œ 78.3์ ์„ ๊ธฐ๋กํ•˜์—ฌ, GPT-40(72.5์ )๋ณด๋‹ค ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋ณด์˜€์Šต๋‹ˆ๋‹ค.
- ๋†’์€ ํ•œ๊ตญ ๋ฌธํ™” ์ดํ•ด๋„: ํ•œ๊ตญ์–ด ๋ฐ ํ•œ๊ตญ ๋ฌธํ™” ๋ฒค์น˜๋งˆํฌ์ธ CLiCk์—์„œ๋„ 83.5์ ์„ ํš๋“ํ•ด, GPT-40(80.2์ )๋ณด๋‹ค ๋” ๋†’์€ ์ดํ•ด๋„๋ฅผ ์ž…์ฆํ–ˆ์Šต๋‹ˆ๋‹ค.
- ํšจ์œจ์ ์ธ ํ† ํฐ ์ฒ˜๋ฆฌ: ๋™์ผํ•œ ํ•œ๊ตญ์–ด ํ…์ŠคํŠธ๋ฅผ ์ž…๋ ฅํ•ด๋„ A.X 4.0๋ณด๋‹ค GPT-40๊ฐ€ ์•ฝ 1.5๋ฐฐ ๋งŽ์€ ํ† ํฐ์„ ์‚ฌ์šฉํ•ฉ๋‹ˆ๋‹ค.
- ๋ฐฉ๋Œ€ํ•œ ์ •๋ณด ์ฒ˜๋ฆฌ: ์ตœ๋Œ€ 131,072 ํ† ํฐ์— ์ด๋ฅด๋Š” ๊ธด ๋ฌธ์„œ๋‚˜ ๋Œ€ํ™”๋„ ํ•œ ๋ฒˆ์— ์ดํ•ดํ•˜๊ณ  ์ฒ˜๋ฆฌํ•  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.
- ๋„๋ฉ”์ธ ์ง€์›: ์ฝ”๋”ฉ, ์ œ์กฐ์—… ๋“ฑ ์ „๋ฌธ ์ง€์‹์ด ํ•„์š”ํ•œ ๋ถ„์•ผ์—์„œ๋„ ํ™œ์šฉํ•  ์ˆ˜ ์žˆ๋„๋ก ๊ธฐ๋ณธ ์„ฑ๋Šฅ์„ ๊ฐ•ํ™”ํ–ˆ์Šต๋‹ˆ๋‹ค.
- ๋ฐฐํฌ ์˜ต์…˜: 720์–ต ๊ฐœ(72B) ๋งค๊ฐœ๋ณ€์ˆ˜๋ฅผ ๊ฐ–์ถ˜ ํ‘œ์ค€ ๋ชจ๋ธ๊ณผ 70์–ต ๊ฐœ(7B) ๋งค๊ฐœ๋ณ€์ˆ˜์˜ ๊ฒฝ๋Ÿ‰ ๋ชจ๋ธ๋กœ ์ œ๊ณต๋˜๋ฉฐ, ๊ธฐ์—… ๋‚ด๋ถ€ ์„œ๋ฒ„์— ์ง์ ‘ ์„ค์น˜(์˜จํ”„๋ ˆ๋ฏธ์Šค)ํ•  ์ˆ˜ ์žˆ์–ด ๋ฐ์ดํ„ฐ ๋ณด์•ˆ์— ๋Œ€ํ•œ ๊ฑฑ์ •์„ ๋œ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

## ํ•ต์‹ฌ ๊ธฐ์ˆ ์€?

### ํ•œ๊ตญ์–ด ํŠนํ™” ํ† ํฌ๋‚˜์ด์ € ์ ์šฉ

ํ•œ๊ตญ์–ด์˜ ๊ณ ์œ ํ•œ ํŠน์„ฑ์„ ์ž˜ ์ดํ•ดํ•˜๋„๋ก ์ตœ์ ํ™”๋œ ํ† ํฌ๋‚˜์ด์ €๋ฅผ ์‚ฌ์šฉํ•ฉ๋‹ˆ๋‹ค. ์ด ํ† ํฌ๋‚˜์ด์ €๋Š” ํ•œ๊ตญ์–ด์˜ ๋‹ค์–‘ํ•œ ํ‘œํ˜„๊ณผ ๋ฌธ๋งฅ์„ ํšจ๊ณผ์ ์œผ๋กœ ํŒŒ์•…ํ•˜๋„๋ก ์„ค๊ณ„๋˜์—ˆ์Šต๋‹ˆ๋‹ค. ๋‚ด๋ถ€ ํ…Œ์ŠคํŠธ ๊ฒฐ๊ณผ, ๊ฐ™์€ ํ•œ๊ตญ์–ด ๋ฌธ์žฅ์„ ์ž…๋ ฅํ–ˆ์„ ๋•Œ GPT-40๋ณด๋‹ค A.X 4.0์ด 33.3% ํšจ์œจ์ ์œผ๋กœ ํ† ํฐ์„ ์‚ฌ์šฉํ•ฉ๋‹ˆ๋‹ค.

์ด๋Š” ์‹ค์ œ ์‚ฌ์šฉ ํ™˜๊ฒฝ์—์„œ ๋‹ค์Œ๊ณผ ๊ฐ™์€ ์žฅ์ ์ด ์žˆ์Šต๋‹ˆ๋‹ค.

- ๊ฐ™์€ ์กฐ๊ฑด์ด๋ผ๋ฉด ๋Œ€๋žต 1.5๋ฐฐ ๋” ๋งŽ์€ ํ•œ๊ตญ์–ด ์ •๋ณด๋ฅผ ์ฒ˜๋ฆฌํ•  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.
- ํ† ํฐ ์ˆ˜๊ฐ€ ์ค„์–ด๋“ค์–ด ์ฒ˜๋ฆฌ ๋น„์šฉ์„ 34% ์ •๋„ ์ ˆ๊ฐํ•  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.
- API๋ฅผ ํ˜ธ์ถœํ•  ๋•Œ ํ† ํฐ ์‚ฌ์šฉ๋Ÿ‰์— ๋”ฐ๋ผ ๋น„์šฉ์ด ์ฑ…์ •๋˜๋Š” ๊ตฌ์กฐ์—์„œ ์œ ๋ฆฌํ•ฉ๋‹ˆ๋‹ค.

ํŠนํžˆ ๋ฌธ์„œ ์š”์•ฝ์ด๋‚˜ ๊ฒ€์ƒ‰ ์ฆ๊ฐ• ์ƒ์„ฑ(RAG) ๋“ฑ ๊ธด ๊ธ€์„ ๋‹ค๋ฃจ๋Š” ๊ธฐ์—… ํ™˜๊ฒฝ์—์„œ, ํ† ํฐ ํšจ์œจ์„ฑ์€ ์šด์˜ ๋น„์šฉ์„ ํฌ๊ฒŒ ์ ˆ๊ฐํ•˜๋Š” ๋ฐ ๊ธฐ์—ฌํ•ฉ๋‹ˆ๋‹ค.

### ํ•œ๊ตญ์–ด ์ดํ•ด์™€ ์ƒ์„ฑ ๋Šฅ๋ ฅ์„ ํ–ฅ์ƒ์‹œํ‚ค๋Š” ํ•™์Šต ๋ฐ์ดํ„ฐ ๊ตฌ์„ฑ

A.X 4.0์— ์‚ฌ์šฉ๋œ ํ•™์Šต ๋ฐ์ดํ„ฐ๋Š” ๋‹ค์Œ๊ณผ ๊ฐ™์€ ํŠน์ง•์„ ๊ฐ–์Šต๋‹ˆ๋‹ค.

- ๊ณ ํ’ˆ์งˆ์˜ ํ•œ๊ตญ์–ด ์ž๋ฃŒ: ์›น์—์„œ ์ถ”์ถœํ•œ ๊ณ ํ’ˆ์งˆ ๋ฐ์ดํ„ฐ, ์ „๋ฌธ ์„œ์ , ํ•ฉ์„ฑ ๋ฐ์ดํ„ฐ๋ฅผ ํฌํ•จํ•œ ๋Œ€๊ทœ๋ชจ ๊ณ ํ’ˆ์งˆ ๋ฐ์ดํ„ฐ์…‹์„ ํ™œ์šฉํ–ˆ์Šต๋‹ˆ๋‹ค.
- ์ฒด๊ณ„์ ์ธ ๋ฐ์ดํ„ฐ ๋ถ„๋ฅ˜: ๋‹ค์–‘ํ•œ ๋ถ„์•ผ์—์„œ ๊ท ํ˜•์žˆ๊ฒŒ ๋†’์€ ์„ฑ๋Šฅ์„ ๋ฐœํœ˜ํ•˜๋„๋ก ์ฃผ์ œ๋ณ„๋กœ ๋ถ„๋ฅ˜๋œ ๋ฐ์ดํ„ฐ์…‹์„ ๊ตฌ์„ฑํ–ˆ์Šต๋‹ˆ๋‹ค.
- ๊ท ํ˜• ์žกํžŒ ์–ธ์–ด ๋ถ„ํฌ: ํ•œ๊ตญ์–ด 42%, ์˜์–ด 51%, ๊ธฐํƒ€ ์–ธ์–ด ๋ฐ ์ฝ”๋“œ 7%๋กœ ๊ตฌ์„ฑํ•ด ์–ธ์–ด ๊ฐ„ ๊ท ํ˜•์„ ์œ ์ง€ํ–ˆ์Šต๋‹ˆ๋‹ค.

์ด๋Ÿฌํ•œ ๋ฐ์ดํ„ฐ ๊ตฌ์„ฑ์€ ๋ชจ๋ธ์ด ํ•œ๊ตญ์–ด์˜ ๋‹ค์–‘ํ•œ ํ‘œํ˜„๊ณผ ๋ฏธ๋ฌ˜ํ•œ ๋ฌธ๋งฅ๊นŒ์ง€ ๊นŠ์ด ์ดํ•ดํ•˜๋„๋ก ๋•์Šต๋‹ˆ๋‹ค.
"""

License

The A.X 4.0 VL Light model is licensed under Apache License 2.0.

Citation

@article{SKTAdotX4VLLight,
  title={A.X 4.0 VL Light},
  author={SKT AI Model Lab},
  year={2025},
  url={https://huggingface.co/skt/A.X-4.0-VL-Light}
}

Contact

Downloads last month
166
Safetensors
Model size
8B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ 2 Ask for provider support

Model tree for skt/A.X-4.0-VL-Light

Finetuned
(15)
this model

Space using skt/A.X-4.0-VL-Light 1

Collection including skt/A.X-4.0-VL-Light