RUIYI

机器视觉智能系统套件

视觉文档与 OCR

集成平台

大多数工厂先完成了系统数字化,纸还没数字化。订单、证书、检验记录和仪表读数仍以纸张、PDF 或某部手机拍下的照片形式到来——然后再由人重新录入一遍。

视觉文档与 OCR 读取页面上的内容:印刷文字、手写体、表格、标记与印章、条码与标签、指针表盘与数显读数。它把这些内容转成命名字段,按你设定的规则校验,再交付给使用这些数据的系统。

「能读」说起来容易,值得信赖却很难。每个结果都带有置信度,低于你设定阈值的结果转人工处理,原始图像始终与它生成的数字关联在一起。

OCR 给你的是字符。你要的是字段。

支持扫描件、照片与 PDF输出字段,而非纯文本置信度与人工复核附带原始图像

它是什么

  • Two kinds of input. Documents on one side — delivery notes, certificates, inspection records, contracts, drawings. Instrument faces and plates on the other — dials, digital displays, nameplates, labels and codes.

  • It reads what is actually on the page. Printed text, handwriting, tables, stamps and barcodes, including pages that arrived skewed, creased, unevenly lit, or photographed at an angle.

  • Output is named fields, not a block of text. Each value comes back with a name, a unit, a type and a confidence value — not something someone has to read through to find what matters.

  • Fields are checked, not just read. Against rules you set: does this order exist, does this quantity exceed what was ordered, is this certificate still in date, is this reading within range.

  • Anything unclear waits for a person. Results below the threshold go to a reviewer with the source image and the reading side by side, rather than being written through.

  • Vision underneath is handled elsewhere. 模型和算法由锐益 Visual CAST 平台提供和管理;本应用关注文档、读数及其后续处理。

当下的阻碍是什么

The paper that never went away. 系统换了,文档没跟上。

  • 供应商单证仍随货流转。送货单、合格证、材质证明、装箱单。记录仍在产线旁、月台旁、巡检路上手工填写。之后由不在场的人录入。读数仍要走到仪表前抄录。把数字抄到一张表上,之后还得再抄一遍。报告仍是纸质的。或是纸质扫描成的 PDF——只是占地更小,问题一模一样。

反复录入。同一份单据进入不止一个系统。

  • 同一批数值被录入不同系统。副本从产生那天起就开始各自漂移。量大时录入慢,量大时也更容易错。两种失效模式结伴而来。数据要等人走完全流程才存在。系统反映的永远是昨天。

识别结果不等于可用数据。一页文字不是一组数值。

  • 识别出的文字仍需人工阅读。才能从整页里挑出真正要紧的几个数值。版式总在变。供应商之间、版本之间、扫描件与拍照之间——固定模板在这三种情况下都会失效。表格跨页、单元格合并、印章压在数字上。而且读数错了也没有任何提示。

Trust and traceability. 数值有了,来源却无从追溯。

  • A figure in the system cannot be traced back to the image it came from. So it cannot be checked without redoing the work.

  • 审核问起数值从何而来,答案只能靠人的记忆。Which is not an answer an audit accepts.

  • Paper gets lost, fades and cannot be searched. Finding one certificate means knowing which box it is in.

  • Nobody knows how much of what was keyed in was ever checked. Because checking leaves no record of its own.

它读取什么

一类是单据文档,另一类是仪表盘面与铭牌标牌。两者的区别在于从图像中提取什么。

输入内容

从中提取什么

送货单、装箱单、收货单据

供应商、零件号、数量、批次、日期

证书与合格文件

标准、牌号、炉批号、检测值、有效期

检验与测试记录、手写日志

记录值、签名、合格与否标记、时间戳

发票、合同与商务单据

金额、日期、当事方、条款、印章

图纸与技术文件

标题栏与明细表、版本、备注

标签、铭牌、条码与矩阵码

序列号、零件号、资产标签、编码标识

仪表面——指针式、数显式与段码屏

读数、单位及其来源仪表

从一张图像到业务系统可用的字段把扫描件、照片和 PDF 识别为文字、表格、手写内容和仪表读数。将结果结构化为命名字段,而非原始文本。按规则和限值校验。低于置信阈值的读数转人工复核,其余继续流转。两条路径最终都把结构化记录送入 ERP、MES、QMS 和档案。读取扫描件、照片、PDF结构化字段,而非文本校验规则与限值交付结构化记录送入 ERP、MES、QMS 和档案复核低置信度

软件产出的是带置信度的读数,不是经过核实的事实。它告诉你字符看起来是什么、把握有多大。阈值由人设定,低于阈值的由人复核,写入记录的内容由人负责。手写体、低对比度印刷、破损页面和光线不佳的仪表都会降低置信度,系统的设计前提就是有些读数需要人介入。原始图像始终附着在每一个字段上,任何数字都能追溯到它读自哪里。

您将获得

纸质单据变成可检索存量档案与日常单据变成可按供应商、零件、批次或日期检索的记录,原始影像一步可达。
是字段,不是文本数值以带单位和类型的命名字段送达,可直接入库,无需通读再录入。
置信度可直接用于行动每个字段都带置信度,低于你设定阈值的会被留下待审,而不是直接写入。
新版式无需发版新单据模板、新仪表类型都能即教即用,无需等待产品变更。
巡检抄表免转录相机读取的读数直接进入记录,誊抄环节及其带来的延迟一并消除。
直入你现有的系统读数、证书与提取的字段直达 ERP、MES、QMS、WMS 或归档系统,而不是滞留在又一个工具里。

应用场景

同一套处理管线,对准不同输入。变化的只是:到的是什么单据、哪些值重要、结果送往何处。

场景

读取什么

收货与供应商单证

送货单、证书与装箱单在入库前完成核验

质量记录与检测报告

检验记录、实验室报告与手写日志录入质量系统

设备巡检与仪表读数

指针、数显与段码屏由相机读取,进入维护与历史记录

工程图纸与技术文件

标题栏、明细表与版本信息用于变更控制与归档

存量档案扫描与归档

历史纸质文件转为可检索记录,原图同步留存

合同、合规与证书管控

关键数值、日期与印章被提取核验,到期与缺失自动标记

核心能力

按功能分组。底层涉及模型、训练与接口的部分,均由锐益 Visual CAST 平台处理。

Capability

What it means

Image pre-processing

在识别之前先完成去噪、倾斜与透视校正、方向检测与纠正,并处理摩尔纹和光照不均。

Printed text recognition

Read printed characters across documents, labels and screens.

Handwriting recognition

Read handwritten entries on forms and logs, with confidence reported per field.

Table recognition

Read tables including merged cells and content that continues across pages.

Layout analysis and field extraction

Locate the parts of a page that matter and take named values from them rather than returning raw text.

Template handling

Apply a template where the page matches one, and fall back to layout analysis where it does not.

Barcode and matrix code reading

Read barcodes and two-dimensional codes on labels, nameplates and documents.

Label and nameplate reading

Read asset tags, serial numbers and rating plates, including worn or low-contrast surfaces.

Instrument reading

Read dial, digital and segmented displays, and record the unit and the instrument the reading came from.

Stamp and mark detection

Detect whether a stamp or seal is present and where, and flag pages where one appears to be missing.

Confidence values

Report confidence per field so downstream steps can decide what needs a person.

Rules and cross-checks

Test readings against reference data — order exists, quantity within tolerance, certificate in date, value within range.

Review queue

Route anything below threshold to a person, showing source image and reading side by side, with the decision recorded.

Searchable output

Produce searchable PDF and structured output so a page can be found by what it says.

Document comparison

Compare two versions of a document and show what changed.

Language coverage

Work across the languages that appear in your documents and your supply chain.

Delivery to business systems

Hand results to ERP, MES, QMS, WMS or archive systems through APIs, files or direct interfaces.

部署方式选择

Run in the cloud, on your own servers, or close to the scanners and cameras — the choice usually depends on where documents may be processed and stored.

Model training and adaptation

用你自己的文档类型和仪表面板训练模型;模型统一通过锐益 Visual CAST 平台管理。

Access and retention

Decide who may submit, view, correct and export, how long images and results are kept, and log every correction against the record it changed.

从图像到可信字段

  1. Capture. 文档通过扫描、上传、邮件或拍照进入系统;仪表示数由固定相机拍摄,或在巡检时用手持设备采集。

  2. Prepare. Each image is cleaned up first — noise removed, orientation and skew corrected, perspective straightened.

  3. Read. Text, tables, handwriting, codes and displays are read, and every value comes back with a confidence figure.

  4. Structure. Readings are mapped to named fields with units and types, using a template where one fits and layout analysis where it does not.

  5. Check. 字段按规则和参考数据校验;超范围、缺失和不一致的字段会被标记。

  6. Review and deliver. Anything below the confidence threshold waits for a person. Everything else is delivered to the systems that use it, with the source image attached to the record.

部署、集成与数据处理

  • Where it runs. In the cloud, on your own servers, or close to the capture device. Documents that may not leave the site, or the country, usually decide this.

  • How it connects. Results are delivered through APIs, files or direct interfaces into ERP, MES, QMS, WMS and archive systems. Models and algorithms underneath are managed by the RUIYI Visual CAST platform.

  • What it starts from. Existing scanners, multifunction devices, mobile cameras and fixed cameras. It does not require replacing them.

  • Who sees what. Submission, viewing, correction and export are controlled by role, and every correction is logged against the record it changed.

  • What is local. 保存期限、含个人信息文档的访问权限以及行业特定规则,取决于所在司法辖区和客户自身政策。系统按这些要求配置;这些决策不由系统做出。