9. data:: 命名空间 — 外部数据
把 PDF 内容与你自己的清单和表格进行核对的 8 个函数。全部在本地处理, 数据不会外传。
9.1 文件的存放位置
术语表和数据集接受相对于执行目录的路径:
data::load_glossary("terms/legal.txt")
data::load_dataset("data/batches.csv")
查询表(query_gtin、query_medicamento、query_postal_code)使用固定
文件名,按以下顺序查找:
$PDFL_DATA_DIR(环境变量)./dados/./- 由
pdfl add安装的配置(pdfl_profiles/*/dados/) - 被分析 PDF 所在目录
PDFL_DATA_DIR=/opt/databases pdfl run profile.pdfl document.pdf
找不到时,错误消息会说明应把文件放在何处。要随配置一起分发,请使用
pdfl pack(第 11 章)。
9.2 术语表与数据集
| 函数 | 功能 |
|---|---|
data::load_glossary(file) |
术语列表(每行一个,# 为注释) |
data::validate_against_reference(file) |
文档中未出现的术语列表 |
data::load_dataset(file) |
把 CSV 读为行的列表 |
data::lookup_value(file, key) |
首列为该键的行的第二列(找不到为 null) |
比较时忽略大小写和空白。
terms/required.txt:
# 每份保单都必须包含的术语
waiting period
covered benefits
general conditions
check "Glossary and dataset" {
terms = data::load_glossary("terms/required.txt")
print("terms in the glossary:", terms.length)
// 最直接的用法
missing = data::validate_against_reference("terms/required.txt")
assert missing.length == 0,
"clauses missing from the policy: #{missing.join("; ")}"
rows = data::load_dataset("data/batches.csv")
print("columns:", rows.first().join(" | ")) // 第一行是表头
print("records:", rows.length - 1)
// null 为假,因此可以直接校验
batch = text::extract_from_region(1, region(400, 50, 150, 20)).trim()
description = data::lookup_value("data/batches.csv", batch)
assert description, "batch #{batch} is not in the approved list"
}
9.3 查询表
按 9.1 的顺序查找固定名称的文件,返回整行列表(找不到则为 null)。
| 函数 | 参照文件 | 功能 |
|---|---|---|
data::query_gtin(code) |
gtin.csv |
按 GTIN 查询(忽略标点) |
data::query_medicamento(reg_or_name) |
medicamentos.csv |
按注册号或名称片段查询 |
data::query_postal_code(code) |
ceps.csv |
按邮编查询(8 位数字) |
data::validate_address(code, "fragment") |
ceps.csv |
该邮编的地址是否包含该片段 |
dados/gtin.csv:
gtin,description,manufacturer
7891234567895,Dipyrone 500mg 20 tablets,Example Labs
check "Lookup tables" {
// 与包装上读取到的条码核对
code = codes::decode_barcode(1)
product = data::query_gtin(code)
assert product, "GTIN #{code} is not in the product database"
print("product:", product.get(2), "| manufacturer:", product.get(3))
// 按注册号查询药品信息
registration = text::extract_from_region(1, region(50, 780, 200, 15)).trim()
medicine = data::query_medicamento(registration)
assert medicine, "registration #{registration} not found"
// 处方药需要法定提示语
band = medicine.get(4)
assert band != "prescription" || text::require_text("PRESCRIPTION ONLY"),
"prescription medicine without the mandatory text"
// 印刷的地址与邮编是否一致
assert data::validate_address("01310100", "Avenida Paulista"),
"printed address does not match the given postal code"
}
9.4 完整示例
// insert_with_databases.pdfl — 与本地数据核对
// 用法: PDFL_DATA_DIR=./databases pdfl run insert_with_databases.pdfl insert.pdf
profile "insert-with-references" {
check "Mandatory regulatory terms" tags: ["glossary"] {
missing = data::validate_against_reference("databases/regulatory_terms.txt")
assert missing.length == 0, "mandatory texts missing: #{missing.join("; ")}"
}
check "Product in the database" tags: ["data", "critical"] {
code = codes::decode_barcode(1)
product = data::query_gtin(code)
assert product, "GTIN #{code} not approved"
// 注册名称必须出现在印刷内容中
name = product.get(2)
assert text::require_text(name),
"the name '#{name}' from the database does not appear on the insert"
print("product verified:", name)
}
check "Registration and band" tags: ["regulatory"] {
registration = text::extract_from_region(1, region(50, 780, 200, 15)).trim()
med = data::query_medicamento(registration)
assert med, "registration #{registration} not found"
assert med.get(4) != "prescription" || text::require_text("PRESCRIPTION ONLY"),
"prescription band requires the prescription notice"
}
check "Manufacturer address" tags: ["data"] {
assert data::validate_address("01310100", "Avenida Paulista"),
"manufacturer address does not match the postal code"
}
}