使用运行时字段探索数据
考虑您想要从中提取字段的大量日志数据。索引这些数据既耗时又占用大量磁盘空间,而您只想在不预先确定 Schema 的情况下探索数据结构。
您知道您的日志数据包含您想要提取的特定字段。在此示例中,我们重点关注 @timestamp 和 message 字段。通过使用运行时字段,您可以定义脚本以在搜索时计算这些字段的值。
您可以从一个简单的示例开始,将 @timestamp 和 message 字段作为索引字段添加到 my-index-000001 映射中。为了保持灵活性,请使用 wildcard 作为 message 的字段类型。
PUT /my-index-000001/
{
"mappings": {
"properties": {
"@timestamp": {
"format": "strict_date_optional_time||epoch_second",
"type": "date"
},
"message": {
"type": "wildcard"
}
}
}
}
映射好您想要检索的字段后,将日志数据中的几条记录索引到 Elasticsearch 中。以下请求使用 bulk API 将原始日志数据索引到 my-index-000001 中。您可以使用一个小样本来试验运行时字段,而不是索引所有日志数据。
最后一个文档不是有效的 Apache 日志格式,但我们可以在脚本中处理这种情况。
POST /my-index-000001/_bulk?refresh
{"index":{}}
{"timestamp":"2020-04-30T14:30:17-05:00","message":"40.135.0.0 - - [30/Apr/2020:14:30:17 -0500] \"GET /images/hm_bg.jpg HTTP/1.0\" 200 24736"}
{"index":{}}
{"timestamp":"2020-04-30T14:30:53-05:00","message":"232.0.0.0 - - [30/Apr/2020:14:30:53 -0500] \"GET /images/hm_bg.jpg HTTP/1.0\" 200 24736"}
{"index":{}}
{"timestamp":"2020-04-30T14:31:12-05:00","message":"26.1.0.0 - - [30/Apr/2020:14:31:12 -0500] \"GET /images/hm_bg.jpg HTTP/1.0\" 200 24736"}
{"index":{}}
{"timestamp":"2020-04-30T14:31:19-05:00","message":"247.37.0.0 - - [30/Apr/2020:14:31:19 -0500] \"GET /french/splash_inet.html HTTP/1.0\" 200 3781"}
{"index":{}}
{"timestamp":"2020-04-30T14:31:22-05:00","message":"247.37.0.0 - - [30/Apr/2020:14:31:22 -0500] \"GET /images/hm_nbg.jpg HTTP/1.0\" 304 0"}
{"index":{}}
{"timestamp":"2020-04-30T14:31:27-05:00","message":"252.0.0.0 - - [30/Apr/2020:14:31:27 -0500] \"GET /images/hm_bg.jpg HTTP/1.0\" 200 24736"}
{"index":{}}
{"timestamp":"2020-04-30T14:31:28-05:00","message":"not a valid apache log"}
此时,您可以查看 Elasticsearch 如何存储您的原始数据。
GET /my-index-000001
该映射包含两个字段:@timestamp 和 message。
{
"my-index-000001" : {
"aliases" : { },
"mappings" : {
"properties" : {
"@timestamp" : {
"type" : "date",
"format" : "strict_date_optional_time||epoch_second"
},
"message" : {
"type" : "wildcard"
},
"timestamp" : {
"type" : "date"
}
}
},
...
}
}
如果您想检索包含 clientip 的结果,可以将该字段作为运行时字段添加到映射中。以下运行时脚本定义了一个 grok 模式,用于从文档中的单个文本字段中提取结构化字段。Grok 模式类似于支持可重用别名表达式的正则表达式。
该脚本匹配 %{{COMMONAPACHELOG}} 日志模式,该模式能够理解 Apache 日志的结构。如果模式匹配(clientip != null),脚本将发出匹配的 IP 地址值。如果模式不匹配,脚本将返回字段值而不会崩溃。
PUT my-index-000001/_mappings
{
"runtime": {
"http.client_ip": {
"type": "ip",
"script": """
String clientip=grok('%{COMMONAPACHELOG}').extract(doc["message"].value)?.clientip;
if (clientip != null) emit(clientip);
"""
}
}
}
- 此条件确保即使消息的模式不匹配,脚本也不会崩溃。
或者,您可以在搜索请求的上下文中定义相同的运行时字段。运行时定义和脚本与之前在索引映射中定义的内容完全相同。将该定义复制到搜索请求的 runtime_mappings 部分下,并包含一个匹配该运行时字段的查询。此查询返回的结果与您在索引映射中为 http.clientip 运行时字段定义搜索查询的结果相同,但仅在此特定搜索的上下文中有效。
GET my-index-000001/_search
{
"runtime_mappings": {
"http.clientip": {
"type": "ip",
"script": """
String clientip=grok('%{COMMONAPACHELOG}').extract(doc["message"].value)?.clientip;
if (clientip != null) emit(clientip);
"""
}
},
"query": {
"match": {
"http.clientip": "40.135.0.0"
}
},
"fields" : ["http.clientip"]
}
您还可以定义一个 composite(复合)运行时字段,以便从单个脚本中发出多个字段。您可以定义一组有类型的子字段并发出一个值映射。在搜索时,每个子字段都会检索映射中与其名称相关联的值。这意味着您只需要指定一次 Grok 模式即可返回多个值。
PUT my-index-000001/_mappings
{
"runtime": {
"http": {
"type": "composite",
"script": "emit(grok(\"%{COMMONAPACHELOG}\").extract(doc[\"message\"].value))",
"fields": {
"clientip": {
"type": "ip"
},
"verb": {
"type": "keyword"
},
"response": {
"type": "long"
}
}
}
}
}
使用 http.clientip 运行时字段,您可以定义一个简单的查询来搜索特定 IP 地址并返回所有相关字段。
GET my-index-000001/_search
{
"query": {
"match": {
"http.clientip": "40.135.0.0"
}
},
"fields" : ["*"]
}
API 返回以下结果。因为 http 是一个 composite 运行时字段,所以响应包含 fields 下的每个子字段,包括匹配查询的任何关联值。无需提前构建数据结构,您就可以以有意义的方式搜索和探索您的数据,从而进行试验并确定要索引哪些字段。
{
...
"hits" : {
"total" : {
"value" : 1,
"relation" : "eq"
},
"max_score" : 1.0,
"hits" : [
{
"_index" : "my-index-000001",
"_id" : "sRVHBnwBB-qjgFni7h_O",
"_score" : 1.0,
"_source" : {
"timestamp" : "2020-04-30T14:30:17-05:00",
"message" : "40.135.0.0 - - [30/Apr/2020:14:30:17 -0500] \"GET /images/hm_bg.jpg HTTP/1.0\" 200 24736"
},
"fields" : {
"http.verb" : [
"GET"
],
"http.clientip" : [
"40.135.0.0"
],
"http.response" : [
200
],
"message" : [
"40.135.0.0 - - [30/Apr/2020:14:30:17 -0500] \"GET /images/hm_bg.jpg HTTP/1.0\" 200 24736"
],
"http.client_ip" : [
"40.135.0.0"
],
"timestamp" : [
"2020-04-30T19:30:17.000Z"
]
}
}
]
}
}
另外,还记得脚本中的那个 if 语句吗?
if (clientip != null) emit(clientip);
如果脚本不包含此条件,查询会在任何不匹配该模式的分片上失败。通过包含此条件,查询会跳过不匹配 Grok 模式的数据。
您还可以运行一个 范围查询 (range query),该查询作用于 timestamp 字段。以下查询返回所有 timestamp 大于或等于 2020-04-30T14:31:27-05:00 的文档。
GET my-index-000001/_search
{
"query": {
"range": {
"timestamp": {
"gte": "2020-04-30T14:31:27-05:00"
}
}
}
}
响应中包括了日志格式不匹配但时间戳在定义范围内的文档。
{
...
"hits" : {
"total" : {
"value" : 2,
"relation" : "eq"
},
"max_score" : 1.0,
"hits" : [
{
"_index" : "my-index-000001",
"_id" : "hdEhyncBRSB6iD-PoBqe",
"_score" : 1.0,
"_source" : {
"timestamp" : "2020-04-30T14:31:27-05:00",
"message" : "252.0.0.0 - - [30/Apr/2020:14:31:27 -0500] \"GET /images/hm_bg.jpg HTTP/1.0\" 200 24736"
}
},
{
"_index" : "my-index-000001",
"_id" : "htEhyncBRSB6iD-PoBqe",
"_score" : 1.0,
"_source" : {
"timestamp" : "2020-04-30T14:31:28-05:00",
"message" : "not a valid apache log"
}
}
]
}
}
如果您不需要正则表达式的强大功能,可以使用 dissect 模式代替 grok 模式。Dissect 模式通过固定分隔符进行匹配,但通常比 grok 更快。
您可以使用 dissect 来实现与使用 grok 模式解析 Apache 日志相同的结果。您不需要匹配日志模式,而是包含您想要丢弃的字符串部分。特别注意您想要丢弃的字符串部分将有助于构建成功的 dissect 模式。
PUT my-index-000001/_mappings
{
"runtime": {
"http.client.ip": {
"type": "ip",
"script": """
String clientip=dissect('%{clientip} %{ident} %{auth} [%{@timestamp}] "%{verb} %{request} HTTP/%{httpversion}" %{status} %{size}').extract(doc["message"].value)?.clientip;
if (clientip != null) emit(clientip);
"""
}
}
}
同样,您可以定义一个 dissect 模式来提取 HTTP 响应代码。
PUT my-index-000001/_mappings
{
"runtime": {
"http.responses": {
"type": "long",
"script": """
String response=dissect('%{clientip} %{ident} %{auth} [%{@timestamp}] "%{verb} %{request} HTTP/%{httpversion}" %{response} %{size}').extract(doc["message"].value)?.response;
if (response != null) emit(Integer.parseInt(response));
"""
}
}
}
然后,您可以使用 http.responses 运行时字段运行查询以检索特定的 HTTP 响应。使用 _search 请求的 fields 参数来指示您想要检索哪些字段。
GET my-index-000001/_search
{
"query": {
"match": {
"http.responses": "304"
}
},
"fields" : ["http.client_ip","timestamp","http.verb"]
}
响应中包括一个 HTTP 响应为 304 的文档。
{
...
"hits" : {
"total" : {
"value" : 1,
"relation" : "eq"
},
"max_score" : 1.0,
"hits" : [
{
"_index" : "my-index-000001",
"_id" : "A2qDy3cBWRMvVAuI7F8M",
"_score" : 1.0,
"_source" : {
"timestamp" : "2020-04-30T14:31:22-05:00",
"message" : "247.37.0.0 - - [30/Apr/2020:14:31:22 -0500] \"GET /images/hm_nbg.jpg HTTP/1.0\" 304 0"
},
"fields" : {
"http.verb" : [
"GET"
],
"http.client_ip" : [
"247.37.0.0"
],
"timestamp" : [
"2020-04-30T19:31:22.000Z"
]
}
}
]
}
}